Fireworks AI
Fireworks AI operates a high-performance AI inference and training platform for open-source models, founded by the former Meta PyTorch engineering team. It serves AI-native startups, enterprises, and developer teams with serverless and dedicated GPU deployments, processing 30+ trillion tokens daily across 10,000+ customers.
- Company typePrivate
- Founded2022
- HeadquartersRedwood City, United States
- Headcount101–250
- GTM typeB2B
- OfferingSoftware
What Fireworks AI does
Fireworks AI is a privately held AI inference and training platform founded in 2022 in Redwood City, California, by the leadership and core engineering team behind PyTorch at Meta — including CEO Lin Qiao (former Head of PyTorch) and CTO Dmytro Dzhulgakov (PyTorch core maintainer). The company operates a disaggregated inference engine optimized from custom kernels through memory management, with proprietary components including FireAttention (claimed 4x faster than vLLM), speculative decoding for sub-100ms streaming responses, multi-node expert parallelism for trillion-parameter MoE models, Multi-LoRA serving, and KV-cache routing for long-context workloads. The training stack spans supervised fine-tuning, DPO, and reinforcement fine-tuning, with a frontier RL toolchain (GRPO, DAPO) and a Training Agent module that takes a model from data upload to production deployment; trained checkpoints deploy to live endpoints in seconds with published K3 KL divergence parity checks. Following the March 2026 acquisition of Hathora, Fireworks runs across 14 regions, multiple bare-metal providers, and four clouds.
The platform serves a hybrid customer base. On the developer side, an API-first self-serve product delivers pay-per-token serverless inference with OpenAI and Anthropic API compatibility, $1 in free credits at signup, cached input tokens at 50% off, and batch inference at half-price; on-demand dedicated GPU deployments are billed per-second at $7–$12 per GPU-hour across H100, H200, B200, and B300 silicon. Enterprise customers buy reserved capacity and Bring-Your-Own-Cloud deployments under multi-year contracts with a 99.9% SLA. The company distributes through Microsoft Foundry (general availability announced at Build 2026), the MongoDB for Startups program, accelerator partnerships with Y Combinator, AWS Activate, and Google for Startups, and a direct field-sales team for Fortune 500 procurement.
Fireworks serves four named customer segments: AI-native startups (Cursor, Vercel, Genspark, Factory AI, Cognition, Lovable), enterprise companies (Samsung, Uber, Doordash, HubSpot, UiPath, GitLab, Notion), developer teams building AI-powered applications (Sourcegraph, Cresta, Quora), and agentic AI builders plus multimodal application developers. Marquee outcomes include Cursor's 13x inference speedup, Notion's latency reduction from 2 seconds to 350 milliseconds, Vercel achieving 90%+ error-free generation at 40x faster throughput, and Factory cutting costs to 6–20% of frontier-model levels. As of May 2026, Fireworks reports processing more than 30 trillion tokens daily for 10,000+ customers at an annualized revenue run-rate of approximately $800 million, with active fundraising discussions at a $15 billion valuation.
Fireworks AI firmographics
Firmographics- Name
- Fireworks AI
- Legal name
- Fireworks AI, Inc.
- Website
- https://fireworks.ai
- Company type
- Private
- Founded year
- 2022
- Operating status
- Operating
- Headcount range
- 101–250 employees
- Short description
- Fireworks AI operates a high-performance AI inference and training platform for open-source models, founded by the former Meta PyTorch engineering team. It serves AI-native startups, enterprises, and developer teams with serverless and dedicated GPU deployments, processing 30+ trillion tokens daily across 10,000+ customers.
- Ownership category
- akta.pro rank
Fireworks AI industry classification
Industry- Product category
- AI Inference Platform
- NAICS
- Software Publishers (5132), Computer Systems Design and Related Services (54151)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management) (HDAEANAA)
- akta.pro secondary industry
- GPU-Accelerated & AI Training/Inference Servers (HDACABAG)
Keywords
Where Fireworks AI is headquartered
LocationHeadquarters
- HQ city
- Redwood City
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Fireworks AI business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Infrastructure, Technology or R&D, Personnel, Marketing or Sales, Operations
Revenue model
- Serverless Inference: Pay-per-token pricing for serverless inference. Cached input tokens priced at 50%. Batch inference also at 50% of serverless pricing. Primary revenue driver for developers and startups.
- On-Demand Deployments: Dedicated GPU deployments with per-GPU-second billing. H100 80GB at $7/hr, H200 141GB at $7/hr, B200 180GB at $10/hr, B300 288GB at $12/hr. No extra charges for startup times.
- Reserved Capacity: Guaranteed capacity with higher quotas, early access to newest regions and hardware. Supports BYOC deployment flexibility. Targets enterprise customers.
- Fine-Tuning Services: LoRA and full-parameter training priced per million training tokens or per GPU hour for reinforcement fine tuning. Base models up to 16B: $0.50-2.00 per 1M tokens depending on method.
- Enterprise Contracts: Direct enterprise sales for reserved capacity and custom deployments. Reported $800M annualized revenue as of May 2026.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | Serverless Inference - Pay per token with zero setup and no cold starts |
| Usage-based | Pay-as-you-go | On-Demand Deployments - Dedicated GPU with per-second billing |
| Usage-based | Pay-as-you-go | Fine-Tuning - Supervised and preference fine tuning |
| Subscription | Multi-year contract | Reserved Capacity - Guaranteed capacity with highest quotas |
| Usage-based | Pay-as-you-go | Embeddings - Text embeddings pricing |
Go-to-market motion3 records
Distribution channels5 records
Marketing channels7 records
Fireworks AI product offering
Product offeringCore offering
Fireworks AI provides a unified AI inference and training platform that runs and fine-tunes 1000+ open-source generative AI models at production scale. The platform offers serverless pay-per-token inference, on-demand dedicated GPU deployments, and reserved enterprise capacity, combined with a Training Platform supporting full-parameter SFT/DPO/RFT. Customers access the service via an OpenAI and Anthropic-compatible API, processing 30+ trillion tokens daily across 14 regions.
Product overview
Fireworks AI is an AI inference and training platform founded by former Meta PyTorch engineers. The platform provides a unified infrastructure for running and fine-tuning open-source AI models, consisting of: (1) Fireworks Inference Platform for serverless, on-demand, and reserved inference with up to 4x throughput gains; (2) Fireworks Training Platform with Training Agent, Managed Training, and Training API for full-spectrum model customization; (3) Model Library with 1000+ open-source models including DeepSeek, Kimi, Qwen, GLM, and NVIDIA Nemotron with day-zero support; (4) Specialized solutions including Code Assistance, Agentic Systems, Multimodal AI, and Enterprise RAG; (5) Acquired Hathora container orchestration platform for real-time workloads. The platform is OpenAI and Anthropic API compatible for easy migration.
Differentiator
Problem solved
Functional benefit
Products and services
- Fireworks Inference Platform High-performance AI inference platform optimized from custom kernels to memory management, delivering up to 4x higher throughput and industry-leading latency. Supports serverless, on-demand, and reserved capacity deployment options for developers and enterprise AI teams running open-source generative models in production.
- Fireworks Training Platform
Quantifiable outcome
- Cursor achieves 13x faster inference speed using Fireworks speculative decoding
- +8 more outcomes
Companies that use Fireworks AI
Customer profileNamed customers25 records
Segments5 records
Ideal customer profiles4 records
Fireworks AI technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration4 records
AI capability10 records
Feature8 records
Fireworks AI partnerships and signals
Strategic signalPartnerships
Eight partnerships are on record, tiered core and minor.
- MicrosoftcoreFireworks AI general availability announced in Microsoft Foundry at Build 2026. Enables enterprise deployment through Azure cloud with unified platform experience. Integration with Microsoft IQ context intelligence layer.
- MongoDB for StartupscoreFireworks AI became launch partner in MongoDB for Startups program expansion in January 2026. Program encompasses companies representing $200B+ combined valuation. Provides matched credit offers across complementary technologies.
- NVIDIAcoreDay-zero availability of NVIDIA Nemotron 3 Super model announced. NVIDIA GTC 2026 featured conversation between Jensen Huang and CEO Lin Qiao. Deep Blackwell platform optimization partnership for 4-10X cost reductions.
- TemporalminorLaunch partner in MongoDB for Startups program alongside Fireworks AI. Provides workflow orchestration complementary to Fireworks inference platform.
- HathoracoreAcquired by Fireworks AI in March 2026. Hathora built global container orchestration platform for latency-sensitive real-time workloads across 14 regions, multiple bare-metal providers, and four clouds. Team joining to enhance compute orchestration for AI inference.
- CursorcoreCursor uses Fireworks as primary inference provider for Composer 2 model. 'Way more performant than open source engines and is what we use in production. RL inference scales elastically and globally because of it.' Fireworks powers Fast Apply and Copilot++ models.
- VercelcoreVercel's v0 model uses Fireworks - 'SOTA changes every day, so you don't want to tie yourself to a single model. Using a fine-tuned RL model with Fireworks, we perform substantially better than Sonnet.' CTO Malte Ubl confirmed partnership.
- UiPathcoreUiPath powers Autopilot and Delegate with open models via Azure Foundry using Fireworks AI. Achieves quality matching Claude Sonnet 4.6 at fraction of cost for Computer Use tasks.
Scale indicators10 records
Recent moves6 records
Expansion highlights7 records
Fireworks AI competitors and assessment
Company assessmentDirect peers
- Modal Labs: Modal offers serverless cloud infrastructure for AI/ML workloads including LLM inference, with a developer-first API and Python-native deployment model. It targets the same API-first, open-model inference use cases as Fireworks.
- OpenRouter: OpenRouter provides a unified API that routes inference requests across multiple model providers and self-hosted backends including Fireworks itself. It is an adjacent routing layer but increasingly competes with Fireworks by aggregating open-model supply.
- Replicate: Replicate runs a cloud for running open-source AI models via API, with a developer-first, pay-per-use model. It overlaps directly with Fireworks' serverless inference and model library offering for open-weight models.
- Baseten: Baseten provides dedicated and serverless inference infrastructure for ML models, with strong emphasis on production deployment, scaling, and performance optimization. It directly competes with Fireworks' on-demand and dedicated deployment offerings.
- Together AI: Together AI offers an open-model inference and fine-tuning cloud very similar to Fireworks, targeting developers and enterprises with GPU clusters and an API for open-weight LLMs. It is the closest direct head-to-head competitor in the open-model inference category.
- Anyscale: Anyscale provides AI compute and inference infrastructure built on Ray, including serverless LLM endpoints for open models. It competes directly with Fireworks on serving open-weight models at scale for AI developers and enterprises.
Broad incumbents
- Hugging Face: Hugging Face hosts open models, datasets, and inference endpoints with a strong developer community and growing enterprise tier (Inference Endpoints, Spaces, and dedicated deployments). It overlaps with Fireworks on open-model serving and developer workflow, with broader scope into model distribution.
- Microsoft Azure AI Foundry: Azure AI Foundry is Microsoft's hyperscaler-scale platform for hosting and serving foundation models, where Fireworks is itself listed as a partner. It competes with Fireworks on enterprise inference procurement while also serving as a distribution channel.
- Google Vertex AI: Vertex AI is Google Cloud's enterprise AI platform covering training, tuning, and inference for both Google and select third-party models. It is a broad incumbent offering that competes with Fireworks for enterprise inference workloads, particularly via Google Cloud's data and MLOps stack.
- AWS Bedrock: AWS Bedrock is a managed service offering foundation models (including open and proprietary) with enterprise-grade security, compliance, and AWS-native billing. It is a broader incumbent alternative to Fireworks for enterprises already standardized on AWS.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks6 records
Key highlights7 records
Customer concentration
Fireworks AI social profiles
Digital presenceFireworks AI financial estimates
Financial estimateRevenue estimate
Valuation estimate
Fireworks AI leadership team
Management profileNumber of profiles
Profiles1 record
Fireworks AI subsidiaries and ownership
Company hierarchySubsidiaries1 record
Fireworks AI funding detail
Funding detailFunding overview
Funding rounds7 records
Investors22 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Fireworks AI M&A and investment
M&A and investmentM&A1 record
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Fireworks AI
What does Fireworks AI do?
Fireworks AI provides a unified AI inference and training platform that runs and fine-tunes 1000+ open-source generative AI models at production scale. The platform offers serverless pay-per-token inference, on-demand dedicated GPU deployments, and reserved enterprise capacity, combined with a Training Platform supporting full-parameter SFT/DPO/RFT. Customers access the service via an OpenAI and Anthropic-compatible API, processing 30+ trillion tokens daily across 14 regions.
Is Fireworks AI a public or private company?
Fireworks AI is a private company. It is classified as venture growth investor backed and is currently operating.
When was Fireworks AI founded?
Fireworks AI was founded in 2022. It employs 101 to 250 people.
Where is Fireworks AI based?
Fireworks AI is headquartered in Redwood City, United States, in the North America region.
How does Fireworks AI make money?
Five revenue lines are on record. Serverless Inference is the primary driver. The others are on-Demand Deployments, reserved Capacity, fine-Tuning Services and enterprise Contracts.
Who are Fireworks AI's main competitors?
Direct peers on record are Modal Labs, OpenRouter, Replicate, Baseten, Together AI and Anyscale. Broad incumbents are Hugging Face, Microsoft Azure AI Foundry, Google Vertex AI and AWS Bedrock.
Does Fireworks AI have an API?
Yes. Fireworks AI offers a serverless inference API that is OpenAI and Anthropic compatible, allowing developers to migrate by swapping a URL. The API supports pay-per-token pricing with serverless deployment, on-demand dedicated deployments, and reserved capacity options. Developers can access models through the Fireworks Python client, REST API, or OpenAI's Python client. Developer documentation is at docs.fireworks.ai/getting-started/introduction.
What industry is Fireworks AI in?
Fireworks AI's product category is AI Inference Platform. Its primary akta.pro industry code is HDAEANAA, End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management), with a secondary code of HDACABAG, GPU-Accelerated & AI Training/Inference Servers. Its NAICS code is 5132 and its SIC code is 7372.