Developer docs
API playgroundTry for free, no card

Search company profiles

Baseten

Full company profile

uuid00000f6

Namestring
Baseten
Legal namestring
Baseten Labs Inc.
Company typeenum
Private
Founded yearint
2019
Descriptiontext

Baseten (Baseten Labs Inc.) is a San Francisco-based AI inference infrastructure company founded in 2019 by Tuhin Srivastava, Amir Haghighat, Phil Howes, and Pankaj Gupta. Its core offering, the Baseten Inference Stack, is a full-stack inference platform combining custom kernels, custom model runtimes, advanced caching, latest decoding techniques, multi-cloud capacity management across roughly 18-20 cloud providers and 87 clusters, and the open-source Truss deployment framework. The platform serves Dedicated Inference deployments for custom and fine-tuned models, pre-optimized Model APIs (per-token pricing), on-demand Training infrastructure, Baseten Chains for compound AI systems, Baseten Embeddings Inference, and the Frontier Gateway white-label product for AI model labs. Customers — including Cursor, Notion, Lovable, Harvey, OpenEvidence, Abridge, Decagon, Gamma, Hebbia, Speechify, Writer, and Zed Industries — span healthcare, legal, developer tools, media, and finance, with disclosed outcomes such as 78% latency reduction at OpenEvidence, 6x cost reduction at Decagon, and 10x+ inference cost reduction at Hebbia.

The business model is primarily usage-based: per-token pricing for Model APIs, per-minute GPU compute for Dedicated Deployments, per-job training compute, and tiered subscription plans (Basic $0/month pay-as-you-go, Pro with volume discounts and dedicated support, Enterprise with custom SLAs, self-hosted deployment, data residency, and multi-year contracts). Distribution combines self-serve product-led growth with enterprise sales and field-deployed engineers, plus channels including Google Cloud Marketplace and white-label OEM via Frontier Gateway. The company was SOC 2 Type II certified and HIPAA compliant as of the disclosure period. Headcount stood at 147 in 2025 with a stated plan to triple in 2026. Capital intensity is high: Baseten had raised approximately $585M total by January 2026 and closed a $1.5B Series F at up to a $13B valuation in June 2026, with NVIDIA investing $150M in the prior Series E. In December 2025 Baseten acquired Parsed, a reinforcement learning startup, to extend capabilities into post-training and continual learning.

Short descriptiontext

Baseten is a San Francisco-based AI inference infrastructure company providing a full-stack inference platform — including dedicated deployments, model APIs, training, and a white-label gateway — for AI-native startups and enterprises to run open-source and custom models at production scale with high reliability and low latency.

Operating statusenum
Operating
Ownership categoryenum
Headcount rangeband
251–500
akta.pro rankint
HeadquartersSan Francisco, United States
HQ citystring
San Francisco
HQ countrystring
United States
HQ regionstring
North America
Markets served

Serves global market

Offices1 record

Each record includes

City, Country, Type, Description, Source

Keyword5 values
AI inference platform, machine learning deployment, model serving infrastructure, GPU inference optimization, cloud inference services
Industry4 codes
1Model Hosting, Serving & Inference Platforms
CodeHDAAACABPrimaryYes
2Model Deployment, Serving & Inference Platforms
CodeHDAAABAFPrimaryNo
3On-Device Inference Runtimes & SDKs (mobile/embedded)
CodeHDAAAJABPrimaryNo
4Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem)
CodeHDAEANACPrimaryNo
NAICS code2 codes
  • Custom Computer Programming Services541511
  • Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services5182
SIC code2 codes
  • Services-Prepackaged Software7372
  • Services-Computer Programming Services7371
Product category
AI Inference Infrastructure
Social media profiles2 records
GTM motion1 record

Each record includes

Type, Description, Source

Revenue model5 records
1Model APIs (Per-Token)
TypeUsage Based
Description

Usage-based pricing per 1M tokens (input, cache input, output) for pre-optimized production models such as GLM 5.2, Kimi K2.7, DeepSeek V4, GPT OSS 120B, NVIDIA Nemotron 3.

baseten.co
2Dedicated Deployments (Per-Minute Compute)
TypeUsage Based
Description

Per-minute compute pricing on dedicated GPU (T4, L4, A10G, A100, H100 MIG, H100, B200) and CPU instances for self-deployed custom/fine-tuned models. Customers only pay for active compute time, not idle.

baseten.co
3Training Compute
TypeUsage Based
Description

On-demand training compute on L4, A10G, A100, H100 MIG, H100, and B200 GPUs for fine-tuning and reinforcement learning jobs.

baseten.co
4Subscription Tiers (Basic / Pro / Enterprise)
TypeSubscription Recurring
Description

Basic plan at $0/month pay-as-you-go; Pro plan adds priority access to high-demand GPUs, dedicated compute, higher Model API rate limits, hands-on engineering, and dedicated Slack/Zoom support; Enterprise plan adds custom SLAs, self-hosted deployments, on-demand flex compute, data residency, advanced security/compliance, custom global regions, and advanced RBAC. Volume discounts available on Pro and Enterprise.

baseten.co
5Frontier Gateway (White-Label Inference APIs)
TypeUsage Based
Description

Recurring, usage-based revenue from managed gateway product enabling AI model labs to monetize their models through white-labeled production APIs with built-in auth, rate limiting, billing, and metering for external billing providers.

tipranks.com
Marketing channels9 records

Each record includes

Title, Type, Stage, Description, Source

Distribution channels6 records

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components5 values
Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations
Pricing details5 tiers
1Basic - $0/month, pay-as-you-go
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Includes dedicated deployments, Model APIs, Training, fast cold starts, SOC 2 Type II and HIPAA compliance, email and in-app chat support; free credits for new accounts; Baseten Cloud deployment.

baseten.co
2Pro - Volume discounts, unlimited autoscaling
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Adds priority access to high-demand GPUs, dedicated compute, higher Model API rate limits, hands-on engineering expertise, dedicated support on Slack and Zoom; Baseten Cloud deployment; quote-based pricing.

baseten.co
3Enterprise - Custom pricing, full control
ModelSubscriptionBilling cadenceMulti-year contract
Notes

Adds custom SLAs, self-host deployments, on-demand flex compute, ability to use existing cloud commitments, full control over data residency, advanced security and compliance, custom global regions, advanced RBAC with Teams; Baseten Cloud, Your VPC, or Hybrid deployment options; volume discounts; quote-based pricing.

baseten.co
4Model APIs - per 1M tokens
ModelUsage-basedBilling cadencePay-as-you-go
Notes

GLM 5.2: $1.50 input cache input / $4.50 output per 1M tokens. GLM 5.1: $1.30 / $4.30. GLM 5: $0.95 / $3.15. GLM 4.7: $0.60 / $2.20. Kimi K2.7 Code/K2.6: $0.95 / $4.00. Kimi K2.5: $0.60 / $3.00. NVIDIA Nemotron 3 Ultra: $0.60 / $2.40. NVIDIA Nemotron 3 Super: $0.30 / $0.75. DeepSeek V4: $1.74 / $3.48. GPT OSS 120B: $0.10 input / $0.50 output.

baseten.co
5Dedicated Deployments - per-minute GPU/CPU compute
ModelUsage-basedBilling cadencePay-as-you-go
Notes

GPU: T4 16GiB $0.01052/min, L4 24GiB $0.01414/min, A10G 24GiB $0.02012/min, A100 80GiB $0.06667/min, H100 MIG 40GiB $0.0625/min, H100 $0.10833/min, B200 180GiB $0.16633/min. CPU: 1x2 $0.00058/min, 1x4 $0.00086/min, 2x8 $0.00173/min, 4x16 $0.00346/min, 8x32 $0.00691/min, 16x64 $0.01382/min.

baseten.co
GTM typeB2B
B2B
Offering typeSoftware
Software
Core offering1 text field

Baseten provides an AI inference platform that enables businesses to deploy, optimize, and scale open-source, custom, and fine-tuned AI models in production. The Baseten Inference Stack combines custom model runtimes, custom kernels, advanced caching, the latest decoding techniques, and multi-cloud capacity management across 18-20 cloud providers and 87 clusters globally. Offerings include Dedicated Deployments for custom models, pre-optimized Model APIs, Training infrastructure, and the Frontier Gateway for white-label inference APIs, supported by Forward Deployed Engineers for hands-on customer support.

Differentiator
Functional benefit
Problem solved
Quantifiable outcome1 of 13 values shown
  • OpenEvidence achieved 78% lower latency (160ms end-to-end), 6x faster deployment, 8x+ reduction in maintenance hours
+12 more records
Product overview1 text field

Baseten offers a unified AI inference platform with a platform-plus-modules architecture centered on the Baseten Inference Stack, which combines the fastest model runtimes, cross-cloud high availability, and seamless developer workflows. The platform encompasses core products Dedicated Inference (for high-scale workloads serving open-source, custom, and fine-tuned models), Model APIs (pre-optimized model endpoints with pay-per-token pricing), Training (for owning model weights with one-click deployment), and Frontier Gateway (managed gateway for AI model labs to deploy white-labeled production inference APIs), supported by deployment options (Baseten Cloud, Self-hosted, and Hybrid), platform modules including Multi-cloud Capacity Management (MCM), Baseten Chains for compound AI, and Baseten Embeddings Inference (BEI), plus the open-source Truss framework and the acquired Parsed post-training capabilities. Together these offerings let customers deploy, optimize, and manage AI models with exceptional latency, reliability, observability, and cost at production scale across more than 15 cloud providers.

Product and service9 records
1Baseten Inference Stack
CategoryAI Inference Platform
Description

The core inference platform that combines the fastest model runtimes, cross-cloud high availability, and seamless developer workflows, powering all Baseten products including Dedicated Deployments, Model APIs, Training, and Frontier Gateway with 99.99% uptime and fast cold starts. Built for businesses deploying production AI workloads at scale.

2Dedicated Deployments
CategoryDedicated Inference
Description

Serves open-source, custom, and fine-tuned AI models on infrastructure purpose-built for high-performance inference at massive scale, with out-of-the-box model performance optimizations and massive horizontal scale. Priced per-minute on dedicated GPU (T4, L4, A10G, A100, H100 MIG, H100, B200) and CPU instances. For businesses deploying custom models with full control and predictable performance.

3Model APIs
CategoryModel APIs
Description

Pre-optimized Model APIs that let users test new workloads, prototype products, or evaluate the latest AI models optimized to be the fastest in production, instantly, with pay-per-token pricing. Includes LLMs (GLM, Kimi, DeepSeek, GPT OSS, NVIDIA Nemotron), transcription (Whisper, Voxtral), TTS (Orpheus, MARS, Qwen3 TTS), and image generation (Flux, Stable Diffusion, Qwen Image, Cosmos).

4Training
CategoryAI Training Infrastructure
Description

Train custom models on Baseten's inference-optimized infrastructure and deploy them in one click, letting customers own their model weights and supporting post-training techniques like SFT, On-Policy Self-Distillation (OPSD), and Iterative SFT. Supports L4, A10G, A100, H100 MIG, H100, and B200 GPUs for supervised fine-tuning and reinforcement learning workflows.

5Frontier Gateway
CategoryWhite-Label API Gateway
Description

A managed gateway for AI model labs to deploy production-grade inference APIs, featuring authentication, authorization, rate limiting, billing integration, and white-label branding with lab-specific URLs and metering for external billing providers. Poolside is a cited customer whose models were converted into whitelabeled production APIs.

6Truss
CategoryOpen-Source ML Framework
Description

Baseten's open-source standard for packaging and serving models built in any framework, enabling developers to deploy any model on Baseten with custom code and dependencies. Hosted on GitHub as a standalone developer tool.

7Baseten Cloud
CategoryDeployment Option
Description

Fully-managed global deployment option with massive horizontal scale, single-tenant clusters for workload isolation, and the fastest time to market for AI inference workloads across 18 cloud providers and 87 clusters.

8Self-hosted Deployment
CategoryDeployment Option
Description

Get the low latency, high throughput, and developer experience of a managed service in your own VPCs, optionally going hybrid with on-demand flex capacity on Baseten Cloud. Full data residency control for regulated enterprises.

9Hybrid Deployment
Scale indicator23 records

Each record includes

Type, Value, Description, Source

Recent move6 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight7 records

Each record includes

Type, Description

Peers10 records
TypeDirect peer
Description

Fireworks AI is a direct inference-platform competitor offering hosted open-source and fine-tuned model serving with usage-based pricing, fast cold starts, and multi-cloud deployment. It targets the same AI-native startup and enterprise customer base as Baseten, and is regularly named alongside Baseten in inference platform comparisons.

TypeDirect peer
Description

Together AI is a direct peer running an inference cloud for open-source LLMs with per-token pricing, dedicated deployments, and training. It competes head-to-head with Baseten for AI startup workloads (chat, code, embeddings) and shares the same open-source, multi-cloud positioning.

TypeDirect peer
Description

Modal provides serverless GPU compute and inference infrastructure for AI workloads with a developer-first, code-driven deployment model. It targets the same AI-native developer audience as Baseten and competes on cold start speed, ease of deployment, and per-second GPU pricing.

TypeDirect peer
Description

Replicate runs a cloud for hosting and serving open-source ML models via API with per-second billing and a large model catalog. It is a direct peer for image, audio, and LLM inference workloads aimed at developers and startups in the same target segments as Baseten.

TypeDirect peer
Description

Anyscale, built on Ray, offers managed compute and inference for AI workloads with a focus on production scaling and custom models. It overlaps with Baseten's Dedicated Deployments for fine-tuned/custom models serving and competes for enterprise customers wanting self-hosted or hybrid Ray-based stacks.

TypeBroad incumbent
Description

AWS SageMaker and Bedrock provide managed model training, deployment, and inference as part of the broader AWS cloud platform. They are a broad incumbent competitor, offering overlapping inference capabilities but bundled into a much wider cloud portfolio rather than specializing in open-source inference the way Baseten does.

TypeBroad incumbent
Description

Google Cloud Vertex AI offers managed training, tuning, and inference for proprietary and open models on Google's TPU/GPU infrastructure. It is a broad incumbent that overlaps with Baseten for enterprise inference, particularly for customers already buying GCP, which is also Baseten's distribution channel via Google Cloud Marketplace.

TypeEmerging player
Description

Groq operates a specialized LPU-based inference cloud emphasizing ultra-low latency for LLM serving. It is an emerging player with partial overlap to Baseten's low-latency inference offering, serving similar customers (voice, code completion, real-time AI) but with proprietary silicon rather than Baseten's GPU-focused, multi-cloud approach.

TypeEmerging player
Description

Cerebras provides inference (and training) services on its proprietary wafer-scale CS systems, marketed for ultra-fast LLM inference. It is an emerging player competing for the same high-performance inference customers as Baseten, but with its own hardware stack rather than multi-cloud GPUs.

TypeOthers
Description

CoreWeave is a large GPU cloud provider supplying NVIDIA H100/B200 capacity that Baseten itself orchestrates through its Multi-cloud Capacity Management. It is an enabling/infrastructure partner and adjacent player rather than a direct inference competitor, though it also offers higher-level inference services for some customers.

Market position
Strengths5 records

Each record includes

Headline, Details, Source

Weaknesses5 records

Each record includes

Headline, Details, Source

Competitive moat6 records

Each record includes

Type, Details

Key risks5 records

Each record includes

Headline, Details, Source

Key highlights7 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Named customers46 records

Each record includes

Name, Industry, Type, Use case, Source, UUID

Segment7 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile5 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
Yes
API detail
Has APIbool
Yes

Docs URL, Description

Integration1 record

Each record includes

Title, Type, Description, Source

AI capability14 records

Each record includes

Type, Description, Source

AI maturity
App detail

Has app

Feature14 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
Profiles6 records

Each record includes

Name, Designation, Designation category, Overview, Profile commentary, Source

Subsidiaries1 record

Each record includes

Name, Acquired on, Relationship type, Type, Business focus

Compliance2 records

Each record includes

Name, Class, Description

Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds9 records

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors28 records

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A2 records

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Baseten

AI Inference Infrastructurebaseten.co

Baseten is a San Francisco-based AI inference infrastructure company providing a full-stack inference platform — including dedicated deployments, model APIs, training, and a white-label gateway — for AI-native startups and enterprises to run open-source and custom models at production scale with high reliability and low latency.

What Baseten does

Baseten (Baseten Labs Inc.) is a San Francisco-based AI inference infrastructure company founded in 2019 by Tuhin Srivastava, Amir Haghighat, Phil Howes, and Pankaj Gupta. Its core offering, the Baseten Inference Stack, is a full-stack inference platform combining custom kernels, custom model runtimes, advanced caching, latest decoding techniques, multi-cloud capacity management across roughly 18-20 cloud providers and 87 clusters, and the open-source Truss deployment framework. The platform serves Dedicated Inference deployments for custom and fine-tuned models, pre-optimized Model APIs (per-token pricing), on-demand Training infrastructure, Baseten Chains for compound AI systems, Baseten Embeddings Inference, and the Frontier Gateway white-label product for AI model labs. Customers — including Cursor, Notion, Lovable, Harvey, OpenEvidence, Abridge, Decagon, Gamma, Hebbia, Speechify, Writer, and Zed Industries — span healthcare, legal, developer tools, media, and finance, with disclosed outcomes such as 78% latency reduction at OpenEvidence, 6x cost reduction at Decagon, and 10x+ inference cost reduction at Hebbia.

The business model is primarily usage-based: per-token pricing for Model APIs, per-minute GPU compute for Dedicated Deployments, per-job training compute, and tiered subscription plans (Basic $0/month pay-as-you-go, Pro with volume discounts and dedicated support, Enterprise with custom SLAs, self-hosted deployment, data residency, and multi-year contracts). Distribution combines self-serve product-led growth with enterprise sales and field-deployed engineers, plus channels including Google Cloud Marketplace and white-label OEM via Frontier Gateway. The company was SOC 2 Type II certified and HIPAA compliant as of the disclosure period. Headcount stood at 147 in 2025 with a stated plan to triple in 2026. Capital intensity is high: Baseten had raised approximately $585M total by January 2026 and closed a $1.5B Series F at up to a $13B valuation in June 2026, with NVIDIA investing $150M in the prior Series E. In December 2025 Baseten acquired Parsed, a reinforcement learning startup, to extend capabilities into post-training and continual learning.

Baseten firmographics

Firmographics
Name
Baseten
Legal name
Baseten Labs Inc.
Website
https://www.baseten.co
Company type
Private
Founded year
2019
Operating status
Operating
Headcount range
251–500 employees
Short description
Baseten is a San Francisco-based AI inference infrastructure company providing a full-stack inference platform — including dedicated deployments, model APIs, training, and a white-label gateway — for AI-native startups and enterprises to run open-source and custom models at production scale with high reliability and low latency.
Ownership category
akta.pro rank

Baseten industry classification

Industry
Product category
AI Inference Infrastructure
NAICS
Custom Computer Programming Services (541511), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
SIC
Services-Prepackaged Software (7372), Services-Computer Programming Services (7371)
akta.pro primary industry
Model Hosting, Serving & Inference Platforms (HDAAACAB)
akta.pro secondary industries
Model Deployment, Serving & Inference Platforms (HDAAABAF), On-Device Inference Runtimes & SDKs (mobile/embedded) (HDAAAJAB), Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC)

Keywords

  • AI inference platform
  • Machine learning deployment
  • Model serving infrastructure
  • GPU inference optimization
  • Cloud inference services

Where Baseten is headquartered

Location

Headquarters

HQ city
San Francisco
HQ country
United States
HQ region
North America

Offices1 record

Markets served

Baseten business model

Business model
GTM type
B2B
Offering type
Software
Cost components
Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations

Revenue model

  1. Model APIs (Per-Token): Usage-based pricing per 1M tokens (input, cache input, output) for pre-optimized production models such as GLM 5.2, Kimi K2.7, DeepSeek V4, GPT OSS 120B, NVIDIA Nemotron 3.
  2. Dedicated Deployments (Per-Minute Compute): Per-minute compute pricing on dedicated GPU (T4, L4, A10G, A100, H100 MIG, H100, B200) and CPU instances for self-deployed custom/fine-tuned models. Customers only pay for active compute time, not idle.
  3. Training Compute: On-demand training compute on L4, A10G, A100, H100 MIG, H100, and B200 GPUs for fine-tuning and reinforcement learning jobs.
  4. Subscription Tiers (Basic / Pro / Enterprise): Basic plan at $0/month pay-as-you-go; Pro plan adds priority access to high-demand GPUs, dedicated compute, higher Model API rate limits, hands-on engineering, and dedicated Slack/Zoom support; Enterprise plan adds custom SLAs, self-hosted deployments, on-demand flex compute, data residency, advanced security/compliance, custom global regions, and advanced RBAC. Volume discounts available on Pro and Enterprise.
  5. Frontier Gateway (White-Label Inference APIs): Recurring, usage-based revenue from managed gateway product enabling AI model labs to monetize their models through white-labeled production APIs with built-in auth, rate limiting, billing, and metering for external billing providers.

Pricing tiers

ModelBillingPrice
Usage-basedPay-as-you-goBasic - $0/month, pay-as-you-go
Usage-basedPay-as-you-goPro - Volume discounts, unlimited autoscaling
SubscriptionMulti-year contractEnterprise - Custom pricing, full control
Usage-basedPay-as-you-goModel APIs - per 1M tokens
Usage-basedPay-as-you-goDedicated Deployments - per-minute GPU/CPU compute

Go-to-market motion1 record

Distribution channels6 records

Marketing channels9 records

Baseten product offering

Product offering

Core offering

Baseten provides an AI inference platform that enables businesses to deploy, optimize, and scale open-source, custom, and fine-tuned AI models in production. The Baseten Inference Stack combines custom model runtimes, custom kernels, advanced caching, the latest decoding techniques, and multi-cloud capacity management across 18-20 cloud providers and 87 clusters globally. Offerings include Dedicated Deployments for custom models, pre-optimized Model APIs, Training infrastructure, and the Frontier Gateway for white-label inference APIs, supported by Forward Deployed Engineers for hands-on customer support.

Product overview

Baseten offers a unified AI inference platform with a platform-plus-modules architecture centered on the Baseten Inference Stack, which combines the fastest model runtimes, cross-cloud high availability, and seamless developer workflows. The platform encompasses core products Dedicated Inference (for high-scale workloads serving open-source, custom, and fine-tuned models), Model APIs (pre-optimized model endpoints with pay-per-token pricing), Training (for owning model weights with one-click deployment), and Frontier Gateway (managed gateway for AI model labs to deploy white-labeled production inference APIs), supported by deployment options (Baseten Cloud, Self-hosted, and Hybrid), platform modules including Multi-cloud Capacity Management (MCM), Baseten Chains for compound AI, and Baseten Embeddings Inference (BEI), plus the open-source Truss framework and the acquired Parsed post-training capabilities. Together these offerings let customers deploy, optimize, and manage AI models with exceptional latency, reliability, observability, and cost at production scale across more than 15 cloud providers.

Differentiator

Problem solved

Functional benefit

Products and services

  • Baseten Inference Stack The core inference platform that combines the fastest model runtimes, cross-cloud high availability, and seamless developer workflows, powering all Baseten products including Dedicated Deployments, Model APIs, Training, and Frontier Gateway with 99.99% uptime and fast cold starts. Built for businesses deploying production AI workloads at scale.
  • Dedicated Deployments Serves open-source, custom, and fine-tuned AI models on infrastructure purpose-built for high-performance inference at massive scale, with out-of-the-box model performance optimizations and massive horizontal scale. Priced per-minute on dedicated GPU (T4, L4, A10G, A100, H100 MIG, H100, B200) and CPU instances. For businesses deploying custom models with full control and predictable performance.
  • Model APIs Pre-optimized Model APIs that let users test new workloads, prototype products, or evaluate the latest AI models optimized to be the fastest in production, instantly, with pay-per-token pricing. Includes LLMs (GLM, Kimi, DeepSeek, GPT OSS, NVIDIA Nemotron), transcription (Whisper, Voxtral), TTS (Orpheus, MARS, Qwen3 TTS), and image generation (Flux, Stable Diffusion, Qwen Image, Cosmos).
  • Training Train custom models on Baseten's inference-optimized infrastructure and deploy them in one click, letting customers own their model weights and supporting post-training techniques like SFT, On-Policy Self-Distillation (OPSD), and Iterative SFT. Supports L4, A10G, A100, H100 MIG, H100, and B200 GPUs for supervised fine-tuning and reinforcement learning workflows.
  • Frontier Gateway A managed gateway for AI model labs to deploy production-grade inference APIs, featuring authentication, authorization, rate limiting, billing integration, and white-label branding with lab-specific URLs and metering for external billing providers. Poolside is a cited customer whose models were converted into whitelabeled production APIs.
  • Truss Baseten's open-source standard for packaging and serving models built in any framework, enabling developers to deploy any model on Baseten with custom code and dependencies. Hosted on GitHub as a standalone developer tool.
  • Baseten Cloud Fully-managed global deployment option with massive horizontal scale, single-tenant clusters for workload isolation, and the fastest time to market for AI inference workloads across 18 cloud providers and 87 clusters.
  • Self-hosted Deployment Get the low latency, high throughput, and developer experience of a managed service in your own VPCs, optionally going hybrid with on-demand flex capacity on Baseten Cloud. Full data residency control for regulated enterprises.
  • Hybrid Deployment

Quantifiable outcome

  • OpenEvidence achieved 78% lower latency (160ms end-to-end), 6x faster deployment, 8x+ reduction in maintenance hours
  • +12 more outcomes

Companies that use Baseten

Customer profile

Named customers46 records

Segments7 records

Ideal customer profiles5 records

Baseten technology and API

Technology

Technology focussed Yes

API detail

Has API
Yes
API docs
API detail

Core technology

AI maturity

App detail

Integration1 record

AI capability14 records

Feature14 records

Baseten partnerships and signals

Strategic signal

Scale indicators23 records

Recent moves6 records

Expansion highlights7 records

Baseten competitors and assessment

Company assessment

Direct peers

  • Fireworks AI: Fireworks AI is a direct inference-platform competitor offering hosted open-source and fine-tuned model serving with usage-based pricing, fast cold starts, and multi-cloud deployment. It targets the same AI-native startup and enterprise customer base as Baseten, and is regularly named alongside Baseten in inference platform comparisons.
  • Together AI: Together AI is a direct peer running an inference cloud for open-source LLMs with per-token pricing, dedicated deployments, and training. It competes head-to-head with Baseten for AI startup workloads (chat, code, embeddings) and shares the same open-source, multi-cloud positioning.
  • Modal: Modal provides serverless GPU compute and inference infrastructure for AI workloads with a developer-first, code-driven deployment model. It targets the same AI-native developer audience as Baseten and competes on cold start speed, ease of deployment, and per-second GPU pricing.
  • Replicate: Replicate runs a cloud for hosting and serving open-source ML models via API with per-second billing and a large model catalog. It is a direct peer for image, audio, and LLM inference workloads aimed at developers and startups in the same target segments as Baseten.
  • Anyscale: Anyscale, built on Ray, offers managed compute and inference for AI workloads with a focus on production scaling and custom models. It overlaps with Baseten's Dedicated Deployments for fine-tuned/custom models serving and competes for enterprise customers wanting self-hosted or hybrid Ray-based stacks.

Broad incumbents

  • AWS (SageMaker / Bedrock): AWS SageMaker and Bedrock provide managed model training, deployment, and inference as part of the broader AWS cloud platform. They are a broad incumbent competitor, offering overlapping inference capabilities but bundled into a much wider cloud portfolio rather than specializing in open-source inference the way Baseten does.
  • Google Cloud Vertex AI: Google Cloud Vertex AI offers managed training, tuning, and inference for proprietary and open models on Google's TPU/GPU infrastructure. It is a broad incumbent that overlaps with Baseten for enterprise inference, particularly for customers already buying GCP, which is also Baseten's distribution channel via Google Cloud Marketplace.

Emerging players

  • Groq: Groq operates a specialized LPU-based inference cloud emphasizing ultra-low latency for LLM serving. It is an emerging player with partial overlap to Baseten's low-latency inference offering, serving similar customers (voice, code completion, real-time AI) but with proprietary silicon rather than Baseten's GPU-focused, multi-cloud approach.
  • Cerebras Systems: Cerebras provides inference (and training) services on its proprietary wafer-scale CS systems, marketed for ultra-fast LLM inference. It is an emerging player competing for the same high-performance inference customers as Baseten, but with its own hardware stack rather than multi-cloud GPUs.

Others

  • CoreWeave: CoreWeave is a large GPU cloud provider supplying NVIDIA H100/B200 capacity that Baseten itself orchestrates through its Multi-cloud Capacity Management. It is an enabling/infrastructure partner and adjacent player rather than a direct inference competitor, though it also offers higher-level inference services for some customers.

Market position

Strengths5 records

Weaknesses5 records

Competitive moat6 records

Key risks5 records

Key highlights7 records

Customer concentration

Baseten social profiles

Digital presence

Baseten compliance and trust

Trust signal

Compliance2 records

Baseten financial estimates

Financial estimate

Revenue estimate

Valuation estimate

Baseten leadership team

Management profile

Number of profiles

Profiles6 records

Baseten subsidiaries and ownership

Company hierarchy

Subsidiaries1 record

Baseten funding detail

Funding detail

Funding overview

Funding rounds9 records

Investors28 records

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

Baseten M&A and investment

M&A and investment

M&A2 records

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about Baseten

What does Baseten do?

Baseten provides an AI inference platform that enables businesses to deploy, optimize, and scale open-source, custom, and fine-tuned AI models in production. The Baseten Inference Stack combines custom model runtimes, custom kernels, advanced caching, the latest decoding techniques, and multi-cloud capacity management across 18-20 cloud providers and 87 clusters globally. Offerings include Dedicated Deployments for custom models, pre-optimized Model APIs, Training infrastructure, and the Frontier Gateway for white-label inference APIs, supported by Forward Deployed Engineers for hands-on customer support.

Is Baseten a public or private company?

Baseten is a private company. It is classified as venture growth investor backed and is currently operating.

When was Baseten founded?

Baseten was founded in 2019. It employs 251 to 500 people.

Where is Baseten based?

Baseten is headquartered in San Francisco, United States, in the North America region.

How does Baseten make money?

Five revenue lines are on record. Model APIs (Per-Token) is the primary driver. The others are dedicated Deployments (Per-Minute Compute), training Compute, subscription Tiers (Basic / Pro / Enterprise) and frontier Gateway (White-Label Inference APIs).

Who are Baseten's main competitors?

Direct peers on record are Fireworks AI, Together AI, Modal, Replicate and Anyscale. Broad incumbents are AWS (SageMaker / Bedrock) and Google Cloud Vertex AI. Emerging players are Groq and Cerebras Systems. CoreWeave is listed as an others.

Does Baseten have an API?

Yes. Baseten offers Model APIs providing instant access to pre-optimized open-source models (LLMs, embeddings, transcription, TTS, image generation) running on the Baseten Inference Stack, priced per 1M tokens. It also exposes a dedicated deployment API for self-serve model deployment. Truss is Baseten's open-source standard for packaging and serving models built in any framework. API endpoints can be white-labeled via Frontier Gateway with authentication, authorization, rate limiting, and billing integration for external billing providers. Developer documentation is at docs.baseten.co.

What industry is Baseten in?

Baseten's product category is AI Inference Infrastructure. Its primary akta.pro industry code is HDAAACAB, Model Hosting, Serving & Inference Platforms, with a secondary code of HDAAABAF, Model Deployment, Serving & Inference Platforms. Its NAICS code is 541511 and its SIC code is 7372.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales
Live signals
SuperpowerdailyBaseten Brings Open Models to OpenAI Enterprise Customers Under Existing CommitmentsBaseten offers task-based routing across open and closed models to OpenAI enterprise customers under existing commitments, with inference on U.S.-based infrastructure and zero prompt retention. The partnership lets customers use Baseten-served open models without shifting spending, but no cost or performance improvements are claimed.RuntimewireBaseten puts Carbon agent sandboxes into preview with NVIDIA OpenShellBaseten put its Carbon agent-execution platform into private preview on September 28th, built on the Blaxel infrastructure it acquired. The preview adds NVIDIA OpenShell runtime for policy enforcement and snapshot-based recovery. The product's commercial viability and production performance remain unproven.SalesTech StarBaseten Names Sheila Vashee Chief Marketing OfficerBaseten appointed Sheila Vashee as Chief Marketing Officer, bringing her experience from Figma and other product-led companies. She will lead marketing as Baseten scales its AI inference infrastructure, which processes 40-50 trillion tokens daily and reported 20X YoY revenue growth in June 2026.citybizBaseten Names Former Figma CMO Sheila Vashee Chief Marketing OfficerBaseten appointed Sheila Vashee chief marketing officer, bringing her three-year Figma CMO tenure where the company surpassed $1 billion in revenue and completed its IPO. She previously led growth at Opendoor and was a second marketing hire at Dropbox. Baseten, an AI inference company, reported 20-fold year-over-year revenue growth in June 2026.YahooBaseten Names Sheila Vashee Chief Marketing OfficerBaseten appointed Sheila Vashee as Chief Marketing Officer, bringing her experience from Figma and other tech companies. She will lead marketing for the AI inference company, which processes 40-50 trillion tokens daily and reported 20X revenue growth in June 2026.MorningstarBaseten Names Sheila Vashee Chief Marketing OfficerBaseten appointed Sheila Vashee as Chief Marketing Officer, bringing her experience from Figma and other tech companies. She joins a leadership team that includes Matt Slagle, Gabe Stern, Sameer Paranjpye, and Vivek Patel. Baseten reported 20X YoY revenue growth in June 2026 and processes 40-50 trillion tokens daily.Analytics InsightTop News Today: Qualcomm’s 2nm AI Chips, Modal-Baseten Funding Talks, Hiring Growth, Tap Global & AI EarbudsQualcomm unveiled two 2nm Snapdragon platforms for AI smartphones, with nine brands including Motorola and OnePlus joining. AI startups Modal and Baseten are in talks for funding at valuations of $15B and $26B, while India's hiring outlook rose to 5.4% for H2 FY27. Tap Global's Tap Earn crypto programme surpassed $8M in assets under management.BloombergStartups Modal, Baseten in Funding Talks to Help Businesses Run AI - BloombergModal Labs and Baseten are in talks to raise capital, with valuations of roughly $15 billion and $26 billion, respectively. The firms provide inference services for running AI models, and spending on inference is expected to eclipse training costs. Both companies declined to comment.BloombergStartups Modal, Baseten in Funding Talks to Help Businesses Run AI - BloombergModal Labs and Baseten are in talks to raise capital, with valuations of roughly $15 billion and $26 billion, respectively. The firms are tapping into the AI inference services market, where spending on chips and computing is expected to eclipse training costs. Both companies declined to comment on the discussions.Crypto BriefingModal, Fireworks, and Baseten gain cost advantage with Nvidia and AMD chips for Kimi K3Modal, Fireworks AI, and Baseten offer hosted inference for Moonshot AI's Kimi K3 model at roughly one-tenth the cost of direct access, using Nvidia and AMD accelerators unavailable to Chinese firms. The 2.8-trillion-parameter model runs at 460 tokens per second, with pricing of $3 per million input tokens and $15 per million output tokens.