Baseten
Baseten is a San Francisco-based AI inference infrastructure company providing a full-stack inference platform — including dedicated deployments, model APIs, training, and a white-label gateway — for AI-native startups and enterprises to run open-source and custom models at production scale with high reliability and low latency.
- Company typePrivate
- Founded2019
- HeadquartersSan Francisco, United States
- Headcount251–500
- GTM typeB2B
- OfferingSoftware
What Baseten does
Baseten (Baseten Labs Inc.) is a San Francisco-based AI inference infrastructure company founded in 2019 by Tuhin Srivastava, Amir Haghighat, Phil Howes, and Pankaj Gupta. Its core offering, the Baseten Inference Stack, is a full-stack inference platform combining custom kernels, custom model runtimes, advanced caching, latest decoding techniques, multi-cloud capacity management across roughly 18-20 cloud providers and 87 clusters, and the open-source Truss deployment framework. The platform serves Dedicated Inference deployments for custom and fine-tuned models, pre-optimized Model APIs (per-token pricing), on-demand Training infrastructure, Baseten Chains for compound AI systems, Baseten Embeddings Inference, and the Frontier Gateway white-label product for AI model labs. Customers — including Cursor, Notion, Lovable, Harvey, OpenEvidence, Abridge, Decagon, Gamma, Hebbia, Speechify, Writer, and Zed Industries — span healthcare, legal, developer tools, media, and finance, with disclosed outcomes such as 78% latency reduction at OpenEvidence, 6x cost reduction at Decagon, and 10x+ inference cost reduction at Hebbia.
The business model is primarily usage-based: per-token pricing for Model APIs, per-minute GPU compute for Dedicated Deployments, per-job training compute, and tiered subscription plans (Basic $0/month pay-as-you-go, Pro with volume discounts and dedicated support, Enterprise with custom SLAs, self-hosted deployment, data residency, and multi-year contracts). Distribution combines self-serve product-led growth with enterprise sales and field-deployed engineers, plus channels including Google Cloud Marketplace and white-label OEM via Frontier Gateway. The company was SOC 2 Type II certified and HIPAA compliant as of the disclosure period. Headcount stood at 147 in 2025 with a stated plan to triple in 2026. Capital intensity is high: Baseten had raised approximately $585M total by January 2026 and closed a $1.5B Series F at up to a $13B valuation in June 2026, with NVIDIA investing $150M in the prior Series E. In December 2025 Baseten acquired Parsed, a reinforcement learning startup, to extend capabilities into post-training and continual learning.
Baseten firmographics
Firmographics- Name
- Baseten
- Legal name
- Baseten Labs Inc.
- Website
- https://www.baseten.co
- Company type
- Private
- Founded year
- 2019
- Operating status
- Operating
- Headcount range
- 251–500 employees
- Short description
- Baseten is a San Francisco-based AI inference infrastructure company providing a full-stack inference platform — including dedicated deployments, model APIs, training, and a white-label gateway — for AI-native startups and enterprises to run open-source and custom models at production scale with high reliability and low latency.
- Ownership category
- akta.pro rank
Baseten industry classification
Industry- Product category
- AI Inference Infrastructure
- NAICS
- Custom Computer Programming Services (541511), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming Services (7371)
- akta.pro primary industry
- Model Hosting, Serving & Inference Platforms (HDAAACAB)
- akta.pro secondary industries
- Model Deployment, Serving & Inference Platforms (HDAAABAF), On-Device Inference Runtimes & SDKs (mobile/embedded) (HDAAAJAB), Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC)
Keywords
Where Baseten is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Baseten business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations
Revenue model
- Model APIs (Per-Token): Usage-based pricing per 1M tokens (input, cache input, output) for pre-optimized production models such as GLM 5.2, Kimi K2.7, DeepSeek V4, GPT OSS 120B, NVIDIA Nemotron 3.
- Dedicated Deployments (Per-Minute Compute): Per-minute compute pricing on dedicated GPU (T4, L4, A10G, A100, H100 MIG, H100, B200) and CPU instances for self-deployed custom/fine-tuned models. Customers only pay for active compute time, not idle.
- Training Compute: On-demand training compute on L4, A10G, A100, H100 MIG, H100, and B200 GPUs for fine-tuning and reinforcement learning jobs.
- Subscription Tiers (Basic / Pro / Enterprise): Basic plan at $0/month pay-as-you-go; Pro plan adds priority access to high-demand GPUs, dedicated compute, higher Model API rate limits, hands-on engineering, and dedicated Slack/Zoom support; Enterprise plan adds custom SLAs, self-hosted deployments, on-demand flex compute, data residency, advanced security/compliance, custom global regions, and advanced RBAC. Volume discounts available on Pro and Enterprise.
- Frontier Gateway (White-Label Inference APIs): Recurring, usage-based revenue from managed gateway product enabling AI model labs to monetize their models through white-labeled production APIs with built-in auth, rate limiting, billing, and metering for external billing providers.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | Basic - $0/month, pay-as-you-go |
| Usage-based | Pay-as-you-go | Pro - Volume discounts, unlimited autoscaling |
| Subscription | Multi-year contract | Enterprise - Custom pricing, full control |
| Usage-based | Pay-as-you-go | Model APIs - per 1M tokens |
| Usage-based | Pay-as-you-go | Dedicated Deployments - per-minute GPU/CPU compute |
Go-to-market motion1 record
Distribution channels6 records
Marketing channels9 records
Baseten product offering
Product offeringCore offering
Baseten provides an AI inference platform that enables businesses to deploy, optimize, and scale open-source, custom, and fine-tuned AI models in production. The Baseten Inference Stack combines custom model runtimes, custom kernels, advanced caching, the latest decoding techniques, and multi-cloud capacity management across 18-20 cloud providers and 87 clusters globally. Offerings include Dedicated Deployments for custom models, pre-optimized Model APIs, Training infrastructure, and the Frontier Gateway for white-label inference APIs, supported by Forward Deployed Engineers for hands-on customer support.
Product overview
Baseten offers a unified AI inference platform with a platform-plus-modules architecture centered on the Baseten Inference Stack, which combines the fastest model runtimes, cross-cloud high availability, and seamless developer workflows. The platform encompasses core products Dedicated Inference (for high-scale workloads serving open-source, custom, and fine-tuned models), Model APIs (pre-optimized model endpoints with pay-per-token pricing), Training (for owning model weights with one-click deployment), and Frontier Gateway (managed gateway for AI model labs to deploy white-labeled production inference APIs), supported by deployment options (Baseten Cloud, Self-hosted, and Hybrid), platform modules including Multi-cloud Capacity Management (MCM), Baseten Chains for compound AI, and Baseten Embeddings Inference (BEI), plus the open-source Truss framework and the acquired Parsed post-training capabilities. Together these offerings let customers deploy, optimize, and manage AI models with exceptional latency, reliability, observability, and cost at production scale across more than 15 cloud providers.
Differentiator
Problem solved
Functional benefit
Products and services
- Baseten Inference Stack The core inference platform that combines the fastest model runtimes, cross-cloud high availability, and seamless developer workflows, powering all Baseten products including Dedicated Deployments, Model APIs, Training, and Frontier Gateway with 99.99% uptime and fast cold starts. Built for businesses deploying production AI workloads at scale.
- Dedicated Deployments Serves open-source, custom, and fine-tuned AI models on infrastructure purpose-built for high-performance inference at massive scale, with out-of-the-box model performance optimizations and massive horizontal scale. Priced per-minute on dedicated GPU (T4, L4, A10G, A100, H100 MIG, H100, B200) and CPU instances. For businesses deploying custom models with full control and predictable performance.
- Model APIs Pre-optimized Model APIs that let users test new workloads, prototype products, or evaluate the latest AI models optimized to be the fastest in production, instantly, with pay-per-token pricing. Includes LLMs (GLM, Kimi, DeepSeek, GPT OSS, NVIDIA Nemotron), transcription (Whisper, Voxtral), TTS (Orpheus, MARS, Qwen3 TTS), and image generation (Flux, Stable Diffusion, Qwen Image, Cosmos).
- Training Train custom models on Baseten's inference-optimized infrastructure and deploy them in one click, letting customers own their model weights and supporting post-training techniques like SFT, On-Policy Self-Distillation (OPSD), and Iterative SFT. Supports L4, A10G, A100, H100 MIG, H100, and B200 GPUs for supervised fine-tuning and reinforcement learning workflows.
- Frontier Gateway A managed gateway for AI model labs to deploy production-grade inference APIs, featuring authentication, authorization, rate limiting, billing integration, and white-label branding with lab-specific URLs and metering for external billing providers. Poolside is a cited customer whose models were converted into whitelabeled production APIs.
- Truss Baseten's open-source standard for packaging and serving models built in any framework, enabling developers to deploy any model on Baseten with custom code and dependencies. Hosted on GitHub as a standalone developer tool.
- Baseten Cloud Fully-managed global deployment option with massive horizontal scale, single-tenant clusters for workload isolation, and the fastest time to market for AI inference workloads across 18 cloud providers and 87 clusters.
- Self-hosted Deployment Get the low latency, high throughput, and developer experience of a managed service in your own VPCs, optionally going hybrid with on-demand flex capacity on Baseten Cloud. Full data residency control for regulated enterprises.
- Hybrid Deployment
Quantifiable outcome
- OpenEvidence achieved 78% lower latency (160ms end-to-end), 6x faster deployment, 8x+ reduction in maintenance hours
- +12 more outcomes
Companies that use Baseten
Customer profileNamed customers46 records
Segments7 records
Ideal customer profiles5 records
Baseten technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration1 record
AI capability14 records
Feature14 records
Baseten partnerships and signals
Strategic signalScale indicators23 records
Recent moves6 records
Expansion highlights7 records
Baseten competitors and assessment
Company assessmentDirect peers
- Fireworks AI: Fireworks AI is a direct inference-platform competitor offering hosted open-source and fine-tuned model serving with usage-based pricing, fast cold starts, and multi-cloud deployment. It targets the same AI-native startup and enterprise customer base as Baseten, and is regularly named alongside Baseten in inference platform comparisons.
- Together AI: Together AI is a direct peer running an inference cloud for open-source LLMs with per-token pricing, dedicated deployments, and training. It competes head-to-head with Baseten for AI startup workloads (chat, code, embeddings) and shares the same open-source, multi-cloud positioning.
- Modal: Modal provides serverless GPU compute and inference infrastructure for AI workloads with a developer-first, code-driven deployment model. It targets the same AI-native developer audience as Baseten and competes on cold start speed, ease of deployment, and per-second GPU pricing.
- Replicate: Replicate runs a cloud for hosting and serving open-source ML models via API with per-second billing and a large model catalog. It is a direct peer for image, audio, and LLM inference workloads aimed at developers and startups in the same target segments as Baseten.
- Anyscale: Anyscale, built on Ray, offers managed compute and inference for AI workloads with a focus on production scaling and custom models. It overlaps with Baseten's Dedicated Deployments for fine-tuned/custom models serving and competes for enterprise customers wanting self-hosted or hybrid Ray-based stacks.
Broad incumbents
- AWS (SageMaker / Bedrock): AWS SageMaker and Bedrock provide managed model training, deployment, and inference as part of the broader AWS cloud platform. They are a broad incumbent competitor, offering overlapping inference capabilities but bundled into a much wider cloud portfolio rather than specializing in open-source inference the way Baseten does.
- Google Cloud Vertex AI: Google Cloud Vertex AI offers managed training, tuning, and inference for proprietary and open models on Google's TPU/GPU infrastructure. It is a broad incumbent that overlaps with Baseten for enterprise inference, particularly for customers already buying GCP, which is also Baseten's distribution channel via Google Cloud Marketplace.
Emerging players
- Groq: Groq operates a specialized LPU-based inference cloud emphasizing ultra-low latency for LLM serving. It is an emerging player with partial overlap to Baseten's low-latency inference offering, serving similar customers (voice, code completion, real-time AI) but with proprietary silicon rather than Baseten's GPU-focused, multi-cloud approach.
- Cerebras Systems: Cerebras provides inference (and training) services on its proprietary wafer-scale CS systems, marketed for ultra-fast LLM inference. It is an emerging player competing for the same high-performance inference customers as Baseten, but with its own hardware stack rather than multi-cloud GPUs.
Others
- CoreWeave: CoreWeave is a large GPU cloud provider supplying NVIDIA H100/B200 capacity that Baseten itself orchestrates through its Multi-cloud Capacity Management. It is an enabling/infrastructure partner and adjacent player rather than a direct inference competitor, though it also offers higher-level inference services for some customers.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks5 records
Key highlights7 records
Customer concentration
Baseten social profiles
Digital presenceBaseten compliance and trust
Trust signalCompliance2 records
Baseten financial estimates
Financial estimateRevenue estimate
Valuation estimate
Baseten leadership team
Management profileNumber of profiles
Profiles6 records
Baseten subsidiaries and ownership
Company hierarchySubsidiaries1 record
Baseten funding detail
Funding detailFunding overview
Funding rounds9 records
Investors28 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Baseten M&A and investment
M&A and investmentM&A2 records
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Baseten
What does Baseten do?
Baseten provides an AI inference platform that enables businesses to deploy, optimize, and scale open-source, custom, and fine-tuned AI models in production. The Baseten Inference Stack combines custom model runtimes, custom kernels, advanced caching, the latest decoding techniques, and multi-cloud capacity management across 18-20 cloud providers and 87 clusters globally. Offerings include Dedicated Deployments for custom models, pre-optimized Model APIs, Training infrastructure, and the Frontier Gateway for white-label inference APIs, supported by Forward Deployed Engineers for hands-on customer support.
Is Baseten a public or private company?
Baseten is a private company. It is classified as venture growth investor backed and is currently operating.
When was Baseten founded?
Baseten was founded in 2019. It employs 251 to 500 people.
Where is Baseten based?
Baseten is headquartered in San Francisco, United States, in the North America region.
How does Baseten make money?
Five revenue lines are on record. Model APIs (Per-Token) is the primary driver. The others are dedicated Deployments (Per-Minute Compute), training Compute, subscription Tiers (Basic / Pro / Enterprise) and frontier Gateway (White-Label Inference APIs).
Who are Baseten's main competitors?
Direct peers on record are Fireworks AI, Together AI, Modal, Replicate and Anyscale. Broad incumbents are AWS (SageMaker / Bedrock) and Google Cloud Vertex AI. Emerging players are Groq and Cerebras Systems. CoreWeave is listed as an others.
Does Baseten have an API?
Yes. Baseten offers Model APIs providing instant access to pre-optimized open-source models (LLMs, embeddings, transcription, TTS, image generation) running on the Baseten Inference Stack, priced per 1M tokens. It also exposes a dedicated deployment API for self-serve model deployment. Truss is Baseten's open-source standard for packaging and serving models built in any framework. API endpoints can be white-labeled via Frontier Gateway with authentication, authorization, rate limiting, and billing integration for external billing providers. Developer documentation is at docs.baseten.co.
What industry is Baseten in?
Baseten's product category is AI Inference Infrastructure. Its primary akta.pro industry code is HDAAACAB, Model Hosting, Serving & Inference Platforms, with a secondary code of HDAAABAF, Model Deployment, Serving & Inference Platforms. Its NAICS code is 541511 and its SIC code is 7372.