Together AI
Together AI operates an AI Native Cloud platform delivering GPU-accelerated compute, serverless and dedicated inference, fine-tuning, and proprietary kernel-level optimizations for 200+ open-source AI models. It serves AI-native startups, generative media companies, and enterprises across 25+ global cities.
- Company typePrivate
- Founded2022
- HeadquartersSan Francisco, United States
- Headcount251–500
- GTM typeB2B
- OfferingSoftware
What Together AI does
Together AI is an AI Native Cloud platform founded in 2022 that provides GPU-accelerated compute infrastructure, inference APIs, fine-tuning, and research-driven optimizations for open-source AI models. The platform supports NVIDIA H100, H200, B200, GB200 NVL72, and GB300 NVL72 GPU clusters with InfiniBand networking, and hosts more than 200 open-source models from providers including DeepSeek, Meta (Llama), Qwen, Google, OpenAI, Mistral, NVIDIA, Moonshot, and MiniMax. The technical differentiator is the Together Kernel Collection—custom CUDA kernels co-developed with Chief Scientist Tri Dao (creator of FlashAttention)—together with ATLAS adaptive speculative decoding, ThunderKittens (an embedded DSL for GPU kernels), Megakernel (single-kernel model forward passes), and FlashAttention-4, which together deliver up to 90% faster training and 2x faster inference versus standard stacks.
The revenue model is built around four usage-based and managed-services streams: per-token pricing on serverless and batch inference APIs (with batch offering up to 50% discount versus real-time), hourly on-demand and reserved GPU cluster compute, dedicated inference deployments with SLA-backed performance, and managed fine-tuning for 100B+ parameter models. Distribution combines API-first product-led growth with a 1M+ developer base and direct enterprise sales for large accounts, with ISO 27001:2022, SOC 2 Type II, and HIPAA-aligned storage available for regulated buyers. The company serves a horizontal customer mix including AI-native startups (Cursor, Cognition, Decagon), generative media platforms (Pika, Hedra, Krea), and enterprises (Salesforce, Zoom, The Washington Post, LG AI Research), reporting 6x to 60x cost savings for customers switching from closed-model APIs.
The company is headquartered in San Francisco with infrastructure deployed across more than 25 cities in the U.S., Europe, and Asia/Middle East, and has raised over $1.3 billion across five funding rounds since 2023. A Series C of $800M led by Aramco Ventures in July 2026 set a post-money valuation of $8.3 billion, with the company reporting annual bookings exceeding $1.15 billion and an annualized revenue run rate of approximately $1 billion as of February 2026. Strategic positioning as an NVIDIA Preferred Partner and participant in the U.S. DOE Genesis Mission underpins both commercial and government-sector reach.
Together AI firmographics
Firmographics- Name
- Together AI
- Legal name
- Together AI
- Website
- https://together.ai
- Company type
- Private
- Founded year
- 2022
- Operating status
- Operating
- Headcount range
- 251–500 employees
- Short description
- Together AI operates an AI Native Cloud platform delivering GPU-accelerated compute, serverless and dedicated inference, fine-tuning, and proprietary kernel-level optimizations for 200+ open-source AI models. It serves AI-native startups, generative media companies, and enterprises across 25+ global cities.
- Ownership category
- akta.pro rank
Together AI industry classification
Industry- Product category
- AI Cloud Infrastructure
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182), Software Publishers (5132), Computer Systems Design and Related Services (54151), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (518)
- SIC
- Services-Computer Programming, Data Processing, Etc. (7370), Services-Prepackaged Software (7372), Services-Computer Integrated Systems Design (7373)
- akta.pro primary industry
- AI Compute Cloud & GPU-as-a-Service (HDAAAAAK)
- akta.pro secondary industries
- Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC), AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers) (HDAAAAAI), End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management) (HDAEANAA), AI Integration & Orchestration Platforms (Connectors, Workflow, iPaaS for AI) (HDAEANAI)
Keywords
Where Together AI is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices12 records
Markets served
Together AI business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Infrastructure, Technology or R&D, Personnel, Marketing or Sales, Operations, Supply Chain
Revenue model
- GPU Cluster Compute (On-Demand & Reserved): Hourly billing for GPU compute capacity on NVIDIA H100, H200, B200, GB200 NVL72, and GB300 NVL72 clusters. On-demand pricing for flexible usage; reserved capacity at lower hourly rates with upfront commitment for 1-6 month terms. Scales from single-node 8-GPU to 4,000+ GPU deployments. Used for model training, fine-tuning, and distributed AI research workloads.
- Serverless Inference API (Per-Token): Pay-per-token pricing for real-time inference APIs across 200+ open-source models including DeepSeek V4 Pro, Llama 4, Qwen3, MiniMax M3, Kimi K2, NVIDIA Nemotron 3, GLM-5, and more. Customers pay based on input and output token volume with no infrastructure management required. Up to 50% cost savings vs. real-time API pricing for batch workloads.
- Batch Inference (Async): Asynchronous batch processing of up to 30 billion enqueued tokens per model per user at up to 50% discount vs. real-time API pricing. Customers upload JSONL files and receive processed results without real-time latency requirements.
- Dedicated Model Inference: Reserved, isolated compute resources for dedicated inference endpoints. Customers deploy specific models on dedicated infrastructure with SLA-backed performance guarantees, providing predictable pricing and consistent latency for production workloads.
- Dedicated Container Inference: GPU infrastructure for running custom containers and generative media models (video, audio, image) with the Sprocket SDK. Supports multi-GPU orchestration and elastic autoscaling. Revenue from container deployment and per-request pricing with no minimum commitments.
- Fine-Tuning Services: Fine-tuning of open-source models using customer data. Supports LoRA and full fine-tuning across model sizes including 100B+ parameter models. Pricing based on training compute time with cost estimation upfront.
- AI Factory (Custom Infrastructure): Custom frontier-scale AI infrastructure deployments for enterprise customers needing 1,000-100,000+ GPU capacity. Reserved long-term capacity with custom configuration for trillion-parameter model training and large-scale inference operations.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Usage-based | Pay-as-you-go | H100 GPU Cluster — On-Demand |
| Usage-based | Multi-year contract | H100 GPU Cluster — Reserved |
| Usage-based | Pay-as-you-go | H200 GPU Cluster — On-Demand |
| Usage-based | Multi-year contract | H200 GPU Cluster — Reserved |
| Usage-based | Pay-as-you-go | B200 GPU Cluster — On-Demand |
| Usage-based | Multi-year contract | B200 GPU Cluster — Reserved |
| Usage-based | Pay-as-you-go | GB200 NVL72 / GB300 NVL72 — Custom |
| Usage-based | Pay-as-you-go | DeepSeek V4 Pro — Serverless Inference |
| Usage-based | Pay-as-you-go | MiniMax M3 — Serverless Inference |
| Usage-based | Pay-as-you-go | NVIDIA Nemotron 3 Ultra — Serverless Inference |
| Usage-based | Pay-as-you-go | GPT-oss-120B — Serverless Inference |
| Usage-based | Pay-as-you-go | Batch Inference — 50% Off Many Top Models |
| Usage-based | Pay-as-you-go | Instant Clusters — Self-Service GPU Provisioning |
| Usage-based | Pay-as-you-go | Batch Inference — Async Processing |
Go-to-market motion2 records
Distribution channels5 records
Marketing channels9 records
Together AI product offering
Product offeringCore offering
Together AI operates an AI Native Cloud platform that provides GPU-accelerated compute, inference APIs, fine-tuning, and proprietary systems research for serving open-source generative AI models at scale. Customers access 200+ open-source models from 40+ providers through serverless, batch, dedicated, and container deployment modes, backed by research-driven kernel optimizations that deliver higher throughput and lower cost than closed-model alternatives.
Product overview
Together AI is an AI Native Cloud platform offering a comprehensive full-stack portfolio of inference, compute, and model shaping products. The core inference layer includes Serverless Inference (auto-scaling pay-per-token API), Batch Inference (asynchronous processing up to 30B tokens), Dedicated Model Inference (reserved GPU infrastructure with proprietary inference engine), and Dedicated Container Inference (for generative media). The accelerated compute layer provides GPU Clusters (self-serve H100/H200/B200/GB200/GB300 with InfiniBand) and Frontier AI Factory (custom 1000+ GPU deployments for trillion-parameter models). Supporting products include Sandbox (fast code environments), Managed Storage (optimized object/parallel storage), and Fine-Tuning (LoRA and full fine-tuning for 100B+ models). Research products including FlashAttention-4, ATLAS, ThunderKittens, Megakernel, and Together Kernel Collection power the platform with memory-efficient attention, adaptive speculative decoding, and custom CUDA kernels achieving up to 4x inference speedups and 90% training acceleration. The platform serves thousands of customers including Cursor, Decagon, Cognition, and The Washington Post.
Differentiator
Problem solved
Functional benefit
Products and services
- Serverless Inference Fully managed pay-per-token inference API with automatic scaling for 200+ open-source models, providing high-performance inference without infrastructure management or long-term commitments. Powers real-time AI workloads for AI-native startups, developers, and enterprise customers.
- Batch Inference Asynchronous batch processing service that scales to 30 billion enqueued tokens per model per user with up to 50% cost savings versus real-time API. Jobs finish within a 24-hour SLA, often within hours, for massive inference workloads.
- Dedicated Model Inference Reserved, isolated compute resources running Together AI's proprietary inference engine (ATLAS adaptive speculative decoding) for production workloads needing consistent latency, control, and best economics with SLA-backed performance guarantees.
- Dedicated Container Inference GPU infrastructure purpose-built for generative media workloads (video, audio, image) running on the Sprocket SDK with multi-GPU orchestration and elastic autoscaling for 10x traffic surges. Multi-cluster scaling handles viral demand for video, audio, and avatar generation models.
- GPU Clusters Self-serve GPU clusters at scale with bare-metal performance, InfiniBand networking, and managed orchestration supporting NVIDIA H100, H200, B200, GB200, and GB300 GPUs. Flexible on-demand pricing ($1.76-$8.19/hr per GPU) and reserved capacity for 1-6 month commitments, scaling from 8 to 4,000+ GPUs.
- Frontier AI Factory Custom infrastructure at frontier scale for trillion-parameter model training and large-scale inference operations, configured for 1,000-100,000+ GPU capacity with NVIDIA Blackwell GPUs, managed orchestration, self-healing infrastructure, and expert support. Reserved long-term capacity for customers including Hedra.
- Fine-Tuning Fine-tuning service for customizing open-source models with user data, supporting LoRA and full fine-tuning for 100B+ parameter models. Includes vision fine-tuning, tool-calling training, and multi-GPU distributed training with upfront cost estimation.
- Sandbox Fast, secure code sandboxes at scale for AI development environments featuring 2.7-second cold starts, 500ms snapshot resumes, VM cloning, and DevContainer support. Powers AI coding workflows including HeroUI's UI component library with 98% lower preview cold starts.
- Managed Storage High-performance managed storage for AI-native workloads providing object storage and parallel filesystems optimized to keep GPUs fed during training and inference. Includes zero egress fees for data movement between regions.
Quantifiable outcome
- Up to 60% lower inference costs vs. proprietary AI APIs, with 6x to 60x savings reported by enterprise customers switching from closed models
- +11 more outcomes
Companies that use Together AI
Customer profileNamed customers20 records
Segments4 records
Ideal customer profiles3 records
Together AI technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration3 records
AI capability12 records
Feature9 records
Together AI partnerships and signals
Strategic signalPartnerships
Seven partnerships are on record, tiered core, strategic and minor.
- PegatroncorePegatron established a strategic collaboration with Together AI and 5C to deliver large-scale AI infrastructure using NVIDIA GB300 NVL72 and HGX B200 liquid-cooled rack deployments in US data centers (Texas and Maryland). Pegatron provides manufacturing and deployment expertise for NVIDIA-powered AI factory infrastructure.
- NVIDIA (Agent Toolkit partner)coreTogether AI listed among NVIDIA Agent Toolkit partners at GTC 2026, alongside CoreWeave, Fireworks, and cloud infrastructure providers integrating NVIDIA's open-source agent development platform.
- Rumble (Quake AI)coreRumble (rebranded as RUM Group Inc.) signed a $270M multi-year GPU cloud agreement in June 2026 to provide dedicated NVIDIA HGX Blackwell B300 GPU capacity to Together AI. Rumble acquired Northern Data's GPU estate (~22,000 NVIDIA H100/H200 GPUs) to fulfill the agreement. This expands Together AI's global GPU footprint with Blackwell-class capacity for large-scale AI training, fine-tuning, and inference workloads.
- AdaptionstrategicTogether AI partnered with Adaption (co-founded by former Cohere and Google DeepMind leaders Sara Hooker and Sudip Roy) to integrate Together Fine-Tuning capabilities natively into Adaption's data management platform. Enables seamless transition from data optimization to model fine-tuning. Adaption reports 82% average increase in data quality for early users.
- U.S. Department of EnergystrategicTogether AI joined the U.S. DOE's Genesis Mission, a national initiative uniting 17 National Laboratories, academia, and private industry to build an integrated AI discovery platform aimed at doubling American scientific productivity within a decade. Together AI contributes FlashAttention and high-performance inference infrastructure to support frontier research in energy and national security.
- CollinearminorCollinear partnered with Together AI to integrate TraitBasis (method for generating realistic simulated users) into the Together Evals platform, enabling builders to test AI models against realistic user behaviors.
- Hugging FacecoreTogether AI integrates with Hugging Face's model hosting ecosystem. Customers can deploy virtually any model from Hugging Face with minimal friction via Dedicated Container Inference (DCI) offering. Together AI listed among top open-source AI model API providers alongside Hugging Face's open-source ecosystem.
Scale indicators21 records
Recent moves8 records
Expansion highlights7 records
Together AI competitors and assessment
Company assessmentDirect peers
- CoreWeave: CoreWeave is a specialized GPU cloud provider offering NVIDIA-powered compute, inference, and AI factory services to AI labs and enterprises. It is the closest direct competitor to Together AI's GPU Clusters and Dedicated Inference offerings and competes head-to-head on H100/H200/B200 capacity.
- Lambda Labs: Lambda operates GPU cloud instances (H100, H200, B200 clusters) and dedicated AI training/inference infrastructure, directly overlapping Together AI's GPU Clusters and AI Factory products for AI-native startups and enterprise research teams.
- Fireworks AI: Fireworks AI provides serverless and dedicated inference APIs for open-source and fine-tuned models with proprietary optimization techniques. It is a direct competitor in the open-model inference API space, targeting the same AI-native developer and enterprise segments as Together AI.
- Replicate: Replicate runs a cloud platform for running and deploying open-source AI models via API, with a focus on generative media (image, video, audio). It overlaps with Together AI's Serverless Inference and Dedicated Container Inference offerings for the same open-model developer ecosystem.
- Anyscale: Anyscale, creator of Ray, offers AI compute and inference platform services on GPU clusters and is repositioning around production AI workloads. It competes with Together AI on enterprise inference, fine-tuning, and large-scale training infrastructure.
- Hugging Face: Hugging Face hosts open-source models and offers Inference API, Endpoints, and Spaces for deploying AI models. It is both a partner and a competitor to Together AI, particularly around open-model hosting and developer-facing inference APIs.
Emerging players
- Modal: Modal provides serverless GPU compute for AI inference and batch workloads, with strong developer ergonomics. It targets a similar developer audience as Together AI's Serverless Inference and Instant Clusters products, though with a smaller model catalog and footprint.
- Nebius: Nebius is rebuilding an AI-focused GPU cloud infrastructure business out of Yandex's hardware estate, offering NVIDIA GPU clusters and managed AI services. It overlaps with Together AI's GPU Clusters and AI Factory offerings, particularly in Europe.
- Crusoe: Crusoe builds large-scale, energy-optimized data centers for AI compute, including NVIDIA GPU clusters and managed AI cloud services. It competes with Together AI in GPU cluster capacity and AI Factory-style deployments for large training runs.
Broad incumbents
- Amazon Web Services (AWS): AWS offers EC2 GPU instances (P5/H100, P6/B200), Bedrock managed inference, and SageMaker AI platforms. As a broad incumbent hyperscaler, AWS competes with Together AI on raw GPU capacity and increasingly on managed inference for both AI-native and enterprise customers.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks6 records
Key highlights7 records
Customer concentration
Together AI social profiles
Digital presenceTogether AI compliance and trust
Trust signalCompliance2 records
Together AI financial estimates
Financial estimateRevenue estimate
Valuation estimate
Together AI leadership team
Management profileNumber of profiles
Profiles8 records
Together AI subsidiaries and ownership
Company hierarchySubsidiaries1 record
Together AI funding detail
Funding detailFunding overview
Funding rounds5 records
Investors41 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Together AI M&A and investment
M&A and investmentM&A2 records
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Together AI
What does Together AI do?
Together AI operates an AI Native Cloud platform that provides GPU-accelerated compute, inference APIs, fine-tuning, and proprietary systems research for serving open-source generative AI models at scale. Customers access 200+ open-source models from 40+ providers through serverless, batch, dedicated, and container deployment modes, backed by research-driven kernel optimizations that deliver higher throughput and lower cost than closed-model alternatives.
Is Together AI a public or private company?
Together AI is a private company. It is classified as venture growth investor backed and is currently operating.
When was Together AI founded?
Together AI was founded in 2022. It employs 251 to 500 people.
Where is Together AI based?
Together AI is headquartered in San Francisco, United States, in the North America region.
How does Together AI make money?
Seven revenue lines are on record. GPU Cluster Compute (On-Demand & Reserved) is the primary driver. The others are serverless Inference API (Per-Token), batch Inference (Async), dedicated Model Inference, dedicated Container Inference, fine-Tuning Services and AI Factory (Custom Infrastructure).
Who are Together AI's main competitors?
Direct peers on record are CoreWeave, Lambda Labs, Fireworks AI, Replicate, Anyscale and Hugging Face. Emerging players are Modal, Nebius and Crusoe. Amazon Web Services (AWS) is listed as a broad incumbent.
Does Together AI have an API?
Yes. Together AI offers a public REST API for inference, fine-tuning, GPU clusters, and model deployment. The API supports serverless inference, batch inference, dedicated model inference, fine-tuning jobs, and cluster management. Authentication via API keys, with endpoints available at api.together.ai. Documentation available at docs.together.ai with SDK support in Python. Developer documentation is at docs.together.ai.
What industry is Together AI in?
Together AI's product category is AI Cloud Infrastructure. Its primary akta.pro industry code is HDAAAAAK, AI Compute Cloud & GPU-as-a-Service, with a secondary code of HDAEANAC, Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem). Its NAICS code is 5182 and its SIC code is 7370.