Developer docs
API playgroundTry for free, no card

Search company profiles

OctoAI

Full company profile

uuid00008qh

Namestring
OctoAI
Legal namestring
OctoAI
Websiteurl
octo.ai
Company typeenum
Private
Founded yearint
2019
Descriptiontext

OctoAI (formerly OctoML) was a Seattle-based AI infrastructure company founded in 2019 as a University of Washington spinout, providing a self-optimizing compute platform for deploying, tuning, and scaling generative AI models. The company's core technology centered on automated inference optimization, with proprietary features including the Asset Orchestrator (enabling dynamic application of thousands of LoRAs and Checkpoints to Stable Diffusion models via a single API) and demonstrated 2.8-second SDXL image generation latency. Its product portfolio spanned developer-facing API services (OctoAI Image Gen) and enterprise-oriented production platforms (OctoStack), with a usage-based revenue model charging customers for AI compute resources consumed.

The company served a hybrid customer base: a primary segment of generative AI developers and startups (including named customers Nightcafe, Storytime AI, and CALA, all of which used image generation workflows), and a secondary enterprise segment targeted through the OctoStack platform. Distribution operated through two channels — self-serve API access for developers globally, and direct field sales for enterprise OctoStack deployments. Customer concentration was dispersed across "dozens of high-growth generative AI customers" with no single anchor customer disclosed.

OctoAI raised approximately $131.9M across four funding rounds from 2019-2021, with the November 2021 Series C ($85M led by Tiger Global Management) valuing the company at $900M post-money. Investors included Addition, Amplify Partners, and Madrona across earlier rounds. In September 2024, NVIDIA acquired OctoAI for approximately $165 million (base) with potential total compensation exceeding $250 million, subsequently winding down independent services by October 31, 2024.

Short descriptiontext

OctoAI (formerly OctoML) was a Seattle-based AI infrastructure platform founded in 2019 that provided self-optimizing compute services for deploying, tuning, and scaling generative AI models, serving developers, startups, and enterprises via API and turn-key platforms. Acquired by NVIDIA in September 2024.

Operating statusenum
Acquired
Ownership categoryenum
Headcount rangeband
51–100
akta.pro rankint
HeadquartersSeattle, United States
HQ citystring
Seattle
HQ countrystring
United States
HQ regionstring
North America
Markets served

Serves global market

Offices1 record

Each record includes

City, Country, Type, Description, Source

Keyword5 values
AI inference platform, generative AI infrastructure, machine learning deployment, AI model optimization, cloud AI compute
Industry3 codes
1Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem)
CodeHDAEANACPrimaryYes
2AI Application Enablement Platforms (Copilot/Agent Frameworks, SDKs)
CodeHDAEANAJPrimaryNo
3End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management)
CodeHDAEANAAPrimaryNo
NAICS code2 codes
  • Software Publishers5132
  • Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services5182
SIC code1 code
  • Services-Prepackaged Software7372
Product category
AI Inference Infrastructure
Social media profiles2 records
GTM motion2 records

Each record includes

Type, Description, Source

Revenue model1 record
1AI Inference Compute Services
TypeUsage Based
Description

OctoAI provided a self-optimizing compute service for AI model deployment and inference. The platform offered flexible hardware options and optimized inference capabilities, operating on a usage-based model for compute resources consumed by AI workloads.

hpcwire.com
Marketing channels2 records

Each record includes

Title, Type, Stage, Description, Source

Distribution channels2 records

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components5 values
Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations
GTM typeB2B
B2B
Offering typeSoftware
Software
Brand1 of 2 records shown
1OctoStack
Description

A turn-key production platform providing optimized inference and model customization for large enterprises.

hpcwire.com
+1 more record
Core offering1 text field

OctoAI provides a self-optimizing AI compute service that helps developers and enterprises run, tune, and scale generative AI models. The platform offers ready-to-use templates for open-source models, flexible hardware options, and an Asset Orchestrator that applies thousands of fine-tuning assets via a single API. It also sells OctoStack, a turn-key production platform for optimized inference and model customization targeting large enterprises.

Differentiator
Functional benefit
Problem solved
Product overview1 text field

OctoAI (formerly OctoML) is an AI compute service company that rebranded in January 2024 to reflect its expanded product suite addressing generative AI market needs. The company offers a self-optimizing compute service with ready-to-use templates for open-source models and flexible hardware options. Its product portfolio includes OctoStack (turn-key production platform for enterprises), OctoAI Image Gen (image generation solution), and an Asset Orchestrator feature. The platform helps developers build and scale AI applications, particularly generative AI models, and was acquired by NVIDIA in September 2024 for approximately $165 million.

Product and service3 records
1OctoAI Self-Optimizing Compute Service
CategoryAI Inference Platform
Description

A self-optimizing compute service that facilitates the development and scaling of AI applications, particularly generative AI models. The platform offers ready-to-use templates for open-source models and flexible hardware options, simplifying AI deployment for developers and enterprises.

2OctoStack
CategoryEnterprise AI Platform
Description

A turn-key production platform providing optimized inference and model customization for large enterprises. OctoStack serves high-growth generative AI customers needing simplified deployment and fine-tuning of AI models in production.

3OctoAI Image Gen
CategoryGenerative AI API
Description

A solution enabling developers to dynamically apply thousands of fine-tuning assets to image generation models built on Stable Diffusion via a single API, delivering image generation in 2.8 seconds on SDXL while eliminating the need to build and maintain custom endpoints.

Scale indicator3 records

Each record includes

Type, Value, Description, Source

Partnership1 partner
Strategic tierCoreTypeTechnology or IntegrationAnnounced on2026-05-13
Description

OctoAI provided a testimonial for AWS EC2 Capacity Blocks for ML, a service that allows customers to reserve accelerated compute instances (GPUs) for machine learning workloads. OctoAI was among multiple companies including NVIDIA, Arcee, Canva, Dashtoon, Leonardo.Ai, and Snorkel that provided positive testimonials about their use cases for training ML models and running AI workloads on this service.

Recent move7 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight4 records

Each record includes

Type, Description

Peers10 records
TypeDirect peer
Description

CoreWeave is a GPU-accelerated cloud provider specializing in large-scale AI workloads, offering optimized inference compute comparable to OctoAI's flexible hardware and optimized inference stack.

TypeDirect peer
Description

Together AI operates an AI inference and fine-tuning cloud platform serving developers and enterprises with optimized open-source model deployment, directly overlapping with OctoAI's self-optimizing compute service and OctoStack offerings.

TypeDirect peer
Description

Anyscale provides a managed Ray-based AI compute platform for training, serving, and scaling ML workloads, competing head-to-head with OctoAI's developer-first API platform for production AI deployment.

TypeDirect peer
Description

Fireworks AI offers an API-first inference platform optimized for open-source generative models with a focus on speed and cost, directly comparable to OctoAI Image Gen and the broader OctoAI inference service.

TypeDirect peer
Description

Replicate runs a cloud API for running and fine-tuning open-source machine learning models, particularly image generation models, sharing the same developer API-first GTM motion and target customer base as OctoAI Image Gen.

TypeDirect peer
Description

Modal provides serverless cloud infrastructure for running AI/ML workloads via an API, targeting developers building production AI applications with a similar ease-of-use and flexible hardware approach to OctoAI.

TypeEmerging player
Description

Hugging Face operates the leading model hub and offers Inference API endpoints for deploying open-source models, overlapping with OctoAI's model serving and fine-tuning capabilities while serving a broader developer community.

TypeDirect peer
Description

RunPod provides GPU cloud infrastructure and serverless inference endpoints for AI workloads, competing with OctoAI on flexible hardware options and developer-friendly deployment of generative AI models.

TypeDirect peer
Description

Lambda operates a GPU cloud and inference platform purpose-built for AI training and serving, sharing the AI compute infrastructure and inference optimization focus that defined OctoAI's offering.

TypeBroad incumbent
Description

AWS offers Bedrock and SageMaker as managed AI services for deploying and serving generative AI models at scale, representing the broad hyperscaler incumbent whose EC2 Capacity Blocks OctoAI explicitly used and endorsed.

Market position
Strengths4 records

Each record includes

Headline, Details, Source

Weaknesses5 records

Each record includes

Headline, Details, Source

Competitive moat3 records

Each record includes

Type, Details

Key risks5 records

Each record includes

Headline, Details, Source

Key highlights6 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Named customers3 records

Each record includes

Name, Industry, Type, Use case, Source, UUID

Segment2 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile2 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
Yes
API detail
Has APIbool
Yes

Docs URL, Description

AI capability3 records

Each record includes

Type, Description, Source

AI maturity
App detail

Has app

Feature3 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
No data
No data
Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds4 records

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors4 records

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

OctoAI

AI Inference Infrastructureocto.ai

OctoAI (formerly OctoML) was a Seattle-based AI infrastructure platform founded in 2019 that provided self-optimizing compute services for deploying, tuning, and scaling generative AI models, serving developers, startups, and enterprises via API and turn-key platforms. Acquired by NVIDIA in September 2024.

What OctoAI does

OctoAI (formerly OctoML) was a Seattle-based AI infrastructure company founded in 2019 as a University of Washington spinout, providing a self-optimizing compute platform for deploying, tuning, and scaling generative AI models. The company's core technology centered on automated inference optimization, with proprietary features including the Asset Orchestrator (enabling dynamic application of thousands of LoRAs and Checkpoints to Stable Diffusion models via a single API) and demonstrated 2.8-second SDXL image generation latency. Its product portfolio spanned developer-facing API services (OctoAI Image Gen) and enterprise-oriented production platforms (OctoStack), with a usage-based revenue model charging customers for AI compute resources consumed.

The company served a hybrid customer base: a primary segment of generative AI developers and startups (including named customers Nightcafe, Storytime AI, and CALA, all of which used image generation workflows), and a secondary enterprise segment targeted through the OctoStack platform. Distribution operated through two channels — self-serve API access for developers globally, and direct field sales for enterprise OctoStack deployments. Customer concentration was dispersed across "dozens of high-growth generative AI customers" with no single anchor customer disclosed.

OctoAI raised approximately $131.9M across four funding rounds from 2019-2021, with the November 2021 Series C ($85M led by Tiger Global Management) valuing the company at $900M post-money. Investors included Addition, Amplify Partners, and Madrona across earlier rounds. In September 2024, NVIDIA acquired OctoAI for approximately $165 million (base) with potential total compensation exceeding $250 million, subsequently winding down independent services by October 31, 2024.

OctoAI firmographics

Firmographics
Name
OctoAI
Legal name
OctoAI
Website
https://octo.ai
Company type
Private
Founded year
2019
Operating status
Acquired
Headcount range
51–100 employees
Short description
OctoAI (formerly OctoML) was a Seattle-based AI infrastructure platform founded in 2019 that provided self-optimizing compute services for deploying, tuning, and scaling generative AI models, serving developers, startups, and enterprises via API and turn-key platforms. Acquired by NVIDIA in September 2024.
Ownership category
akta.pro rank

OctoAI industry classification

Industry
Product category
AI Inference Infrastructure
NAICS
Software Publishers (5132), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
SIC
Services-Prepackaged Software (7372)
akta.pro primary industry
Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC)
akta.pro secondary industries
AI Application Enablement Platforms (Copilot/Agent Frameworks, SDKs) (HDAEANAJ), End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management) (HDAEANAA)

Keywords

  • AI inference platform
  • Generative AI infrastructure
  • Machine learning deployment
  • AI model optimization
  • Cloud AI compute

Where OctoAI is headquartered

Location

Headquarters

HQ city
Seattle
HQ country
United States
HQ region
North America

Offices1 record

Markets served

OctoAI business model

Business model
GTM type
B2B
Offering type
Software
Cost components
Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations

Revenue model

  1. AI Inference Compute Services: OctoAI provided a self-optimizing compute service for AI model deployment and inference. The platform offered flexible hardware options and optimized inference capabilities, operating on a usage-based model for compute resources consumed by AI workloads.

Go-to-market motion2 records

Distribution channels2 records

Marketing channels2 records

OctoAI product offering

Product offering

Core offering

OctoAI provides a self-optimizing AI compute service that helps developers and enterprises run, tune, and scale generative AI models. The platform offers ready-to-use templates for open-source models, flexible hardware options, and an Asset Orchestrator that applies thousands of fine-tuning assets via a single API. It also sells OctoStack, a turn-key production platform for optimized inference and model customization targeting large enterprises.

Product overview

OctoAI (formerly OctoML) is an AI compute service company that rebranded in January 2024 to reflect its expanded product suite addressing generative AI market needs. The company offers a self-optimizing compute service with ready-to-use templates for open-source models and flexible hardware options. Its product portfolio includes OctoStack (turn-key production platform for enterprises), OctoAI Image Gen (image generation solution), and an Asset Orchestrator feature. The platform helps developers build and scale AI applications, particularly generative AI models, and was acquired by NVIDIA in September 2024 for approximately $165 million.

Differentiator

Problem solved

Functional benefit

Brands

  • OctoStack: A turn-key production platform providing optimized inference and model customization for large enterprises.
  • OctoAI Image Gen

Products and services

  • OctoAI Self-Optimizing Compute Service A self-optimizing compute service that facilitates the development and scaling of AI applications, particularly generative AI models. The platform offers ready-to-use templates for open-source models and flexible hardware options, simplifying AI deployment for developers and enterprises.
  • OctoStack A turn-key production platform providing optimized inference and model customization for large enterprises. OctoStack serves high-growth generative AI customers needing simplified deployment and fine-tuning of AI models in production.
  • OctoAI Image Gen A solution enabling developers to dynamically apply thousands of fine-tuning assets to image generation models built on Stable Diffusion via a single API, delivering image generation in 2.8 seconds on SDXL while eliminating the need to build and maintain custom endpoints.

Companies that use OctoAI

Customer profile

Named customers3 records

Segments2 records

Ideal customer profiles2 records

OctoAI technology and API

Technology

Technology focussed Yes

API detail

Has API
Yes
API docs
API detail

Core technology

AI maturity

App detail

AI capability3 records

Feature3 records

OctoAI partnerships and signals

Strategic signal

Partnerships

One partnership is on record.

  • Amazon Web Services (AWS)coreTechnology or Integration · 13 May 2026OctoAI provided a testimonial for AWS EC2 Capacity Blocks for ML, a service that allows customers to reserve accelerated compute instances (GPUs) for machine learning workloads. OctoAI was among multiple companies including NVIDIA, Arcee, Canva, Dashtoon, Leonardo.Ai, and Snorkel that provided positive testimonials about their use cases for training ML models and running AI workloads on this service.

Scale indicators3 records

Recent moves7 records

Expansion highlights4 records

OctoAI competitors and assessment

Company assessment

Direct peers

  • CoreWeave: CoreWeave is a GPU-accelerated cloud provider specializing in large-scale AI workloads, offering optimized inference compute comparable to OctoAI's flexible hardware and optimized inference stack.
  • Together AI: Together AI operates an AI inference and fine-tuning cloud platform serving developers and enterprises with optimized open-source model deployment, directly overlapping with OctoAI's self-optimizing compute service and OctoStack offerings.
  • Anyscale: Anyscale provides a managed Ray-based AI compute platform for training, serving, and scaling ML workloads, competing head-to-head with OctoAI's developer-first API platform for production AI deployment.
  • Fireworks AI: Fireworks AI offers an API-first inference platform optimized for open-source generative models with a focus on speed and cost, directly comparable to OctoAI Image Gen and the broader OctoAI inference service.
  • Replicate: Replicate runs a cloud API for running and fine-tuning open-source machine learning models, particularly image generation models, sharing the same developer API-first GTM motion and target customer base as OctoAI Image Gen.
  • Modal Labs: Modal provides serverless cloud infrastructure for running AI/ML workloads via an API, targeting developers building production AI applications with a similar ease-of-use and flexible hardware approach to OctoAI.
  • RunPod: RunPod provides GPU cloud infrastructure and serverless inference endpoints for AI workloads, competing with OctoAI on flexible hardware options and developer-friendly deployment of generative AI models.
  • Lambda Labs: Lambda operates a GPU cloud and inference platform purpose-built for AI training and serving, sharing the AI compute infrastructure and inference optimization focus that defined OctoAI's offering.

Emerging players

  • Hugging Face: Hugging Face operates the leading model hub and offers Inference API endpoints for deploying open-source models, overlapping with OctoAI's model serving and fine-tuning capabilities while serving a broader developer community.

Broad incumbents

  • Amazon Web Services (Bedrock / SageMaker): AWS offers Bedrock and SageMaker as managed AI services for deploying and serving generative AI models at scale, representing the broad hyperscaler incumbent whose EC2 Capacity Blocks OctoAI explicitly used and endorsed.

Market position

Strengths4 records

Weaknesses5 records

Competitive moat3 records

Key risks5 records

Key highlights6 records

Customer concentration

OctoAI social profiles

Digital presence

OctoAI financial estimates

Financial estimate

Revenue estimate

Valuation estimate

OctoAI leadership team

Management profile

Number of profiles

OctoAI funding detail

Funding detail

Funding overview

Funding rounds4 records

Investors4 records

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

OctoAI M&A and investment

M&A and investment

M&A

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about OctoAI

What does OctoAI do?

OctoAI provides a self-optimizing AI compute service that helps developers and enterprises run, tune, and scale generative AI models. The platform offers ready-to-use templates for open-source models, flexible hardware options, and an Asset Orchestrator that applies thousands of fine-tuning assets via a single API. It also sells OctoStack, a turn-key production platform for optimized inference and model customization targeting large enterprises.

Is OctoAI a public or private company?

OctoAI is a private company. It is classified as corporate owned and is currently acquired.

When was OctoAI founded?

OctoAI was founded in 2019. It employs 51 to 100 people.

Where is OctoAI based?

OctoAI is headquartered in Seattle, United States, in the North America region.

How does OctoAI make money?

One revenue line is on record: AI Inference Compute Services.

Who are OctoAI's main competitors?

Direct peers on record are CoreWeave, Together AI, Anyscale, Fireworks AI, Replicate, Modal Labs, RunPod and Lambda Labs. Hugging Face is listed as an emerging player. Amazon Web Services (Bedrock / SageMaker) is listed as a broad incumbent.

Does OctoAI have an API?

Yes. OctoAI offers an API that enables developers to dynamically apply thousands of fine-tuning assets to image generation models built on Stable Diffusion via a single API, delivering image generation in 2.8 seconds on SDXL.

What industry is OctoAI in?

OctoAI's product category is AI Inference Infrastructure. Its primary akta.pro industry code is HDAEANAC, Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem), with a secondary code of HDAEANAJ, AI Application Enablement Platforms (Copilot/Agent Frameworks, SDKs). Its NAICS code is 5132 and its SIC code is 7372.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales
Live signals
Amazon Web ServicesAmazon EC2 Capacity Blocks for MLAmazon Web Services (AWS) announced and detailed EC2 Capacity Blocks for ML, a service that allows customers to reserve accelerated compute instances (GPUs) in Amazon EC2 UltraClusters for machine learning workloads, with capacity available up to eight weeks in advance and clusters ranging from one to 64 instances (up to 512 GPUs). The service supports NVIDIA GPUs (H100, H200, A100, Blackwell-based), AWS Trainium chips, and instances can be reserved for up to six months. Multiple companies including NVIDIA, Arcee, Canva, Dashtoon, Leonardo.Ai, OctoAI, and Snorkel provided testimonials about their use cases and positive experiences with the service for training ML models and running AI workloads.AI CERTsHow Model Compression Platforms Power Edge AI CommercializationThe model compression platform market is experiencing rapid growth, with Grand View Research projecting the global edge-AI market to reach USD 24.9 billion in 2025 and expand at a 21.7 percent CAGR through 2033, fueling increased venture investment and strategic acquisitions. Major technology firms are aggressively acquiring compression tooling providers: NVIDIA purchased OctoML (now OctoAI) for an estimated USD 165–250 million in September 2024, Red Hat acquired Neural Magic in January 2025, and Qualcomm announced plans to integrate Edge Impulse with its Dragonwing processors in March 2025, consolidating optimization capabilities around their respective silicon ecosystems. Industry analysts warn that rapid consolidation could deepen vendor lock-in as proprietary runtimes displace open standards, prompting enterprises to demand neutral model compression platforms that balance flexibility with performance gains of 30–60 percent total inference cost savings.StartupSeekerStartupSeekerOctoAI, an AI infrastructure platform provider, has been acquired by NVIDIA. The platform enables developers to efficiently run, tune, and scale generative AI models with features like low-latency inference and enterprise-grade reliability.CyberdboctoAINVIDIA Corporation acquired octoAI in September 2024, resulting in the dissolution of octoAI as an independent corporate entity. The acquisition integrates octoAI's generative AI infrastructure platform into NVIDIA's operations.BeehiivNvidia Acquires OctoAI, ChatGPT Prompts, AI Job MarketNvidia has acquired Seattle-based startup OctoAI in a $250 million deal to strengthen its enterprise generative AI infrastructure. This transaction marks Nvidia’s fifth major acquisition in 2024, aiming to make enterprise-level AI more accessible and scalable for businesses.WebcatalogDesktop App for Mac, Windows (PC)OctoAI provides infrastructure for running, tuning, and scaling generative AI applications to enable developers to integrate powerful models. A desktop application is available via WebCatalog for macOS and Windows to facilitate easier account management and multitasking.HpcwireData Science • AI • Advanced AnalyticsNvidia has acquired OctoAI, a Seattle-based AI infrastructure startup known for its Apache TVM abstraction layer, in a deal valued at approximately $165 million to potentially over $250 million. Following the acquisition, OctoAI is winding down its commercial services and terminating customer access effective October 31, 2024, while its founder and CEO Luis Ceze joins Nvidia as Vice President of AI systems software.KisacoresearchNvidia Acquires OctoAI To Dominate Enterprise Generative AI SolutionsNvidia has acquired Seattle-based OctoAI, a startup focused on generative AI tools, in a $250 million deal reported by Forbes. The purchase is Nvidia's fifth acquisition of 2024 and adds OctoAI's hardware-agnostic cloud platform, which supports multiple chip architectures including AMD and Intel. OctoAI had planned vertical-specific offerings, including healthcare.AutoizeOctoAI Acquired by NVIDIA – AI Inference & GenAI AlternativesA blog post reports that OctoAI was acquired by NVIDIA for a reported $165 million, with the OctoAI team being acquired and its OctoStack IP absorbed to improve NVIDIA NIM inference performance. OctoAI will wind down its commercial services, terminating customer access on October 31, 2024, and merging into NVIDIA.HpcwireData Science • AI • Advanced AnalyticsOctoAI, formerly known as OctoML, rebranded in January to reflect its expanded product suite addressing generative AI market needs, offering platforms for developers to build production applications with various AI models. The company launched OctoStack, a turn-key production platform providing optimized inference and model customization for large enterprises, and serves dozens of high-growth generative AI customers. The co-founder discusses 2024 as the year generative AI transitions from experimentation to production, emphasizing the importance of controlling inference costs, selecting the right model, and leveraging fine-tuning techniques for customization.