Developer docs
API playgroundTry for free, no card

Search company profiles

Together AI

Full company profile

uuid00000f2

Namestring
Together AI
Legal namestring
Together AI
Websiteurl
together.ai
Company typeenum
Private
Founded yearint
2022
Descriptiontext

Together AI is an AI Native Cloud platform founded in 2022 that provides GPU-accelerated compute infrastructure, inference APIs, fine-tuning, and research-driven optimizations for open-source AI models. The platform supports NVIDIA H100, H200, B200, GB200 NVL72, and GB300 NVL72 GPU clusters with InfiniBand networking, and hosts more than 200 open-source models from providers including DeepSeek, Meta (Llama), Qwen, Google, OpenAI, Mistral, NVIDIA, Moonshot, and MiniMax. The technical differentiator is the Together Kernel Collection—custom CUDA kernels co-developed with Chief Scientist Tri Dao (creator of FlashAttention)—together with ATLAS adaptive speculative decoding, ThunderKittens (an embedded DSL for GPU kernels), Megakernel (single-kernel model forward passes), and FlashAttention-4, which together deliver up to 90% faster training and 2x faster inference versus standard stacks.

The revenue model is built around four usage-based and managed-services streams: per-token pricing on serverless and batch inference APIs (with batch offering up to 50% discount versus real-time), hourly on-demand and reserved GPU cluster compute, dedicated inference deployments with SLA-backed performance, and managed fine-tuning for 100B+ parameter models. Distribution combines API-first product-led growth with a 1M+ developer base and direct enterprise sales for large accounts, with ISO 27001:2022, SOC 2 Type II, and HIPAA-aligned storage available for regulated buyers. The company serves a horizontal customer mix including AI-native startups (Cursor, Cognition, Decagon), generative media platforms (Pika, Hedra, Krea), and enterprises (Salesforce, Zoom, The Washington Post, LG AI Research), reporting 6x to 60x cost savings for customers switching from closed-model APIs.

The company is headquartered in San Francisco with infrastructure deployed across more than 25 cities in the U.S., Europe, and Asia/Middle East, and has raised over $1.3 billion across five funding rounds since 2023. A Series C of $800M led by Aramco Ventures in July 2026 set a post-money valuation of $8.3 billion, with the company reporting annual bookings exceeding $1.15 billion and an annualized revenue run rate of approximately $1 billion as of February 2026. Strategic positioning as an NVIDIA Preferred Partner and participant in the U.S. DOE Genesis Mission underpins both commercial and government-sector reach.

Short descriptiontext

Together AI operates an AI Native Cloud platform delivering GPU-accelerated compute, serverless and dedicated inference, fine-tuning, and proprietary kernel-level optimizations for 200+ open-source AI models. It serves AI-native startups, generative media companies, and enterprises across 25+ global cities.

Operating statusenum
Operating
Ownership categoryenum
Headcount rangeband
251–500
akta.pro rankint
HeadquartersSan Francisco, United States
HQ citystring
San Francisco
HQ countrystring
United States
HQ regionstring
North America
Markets served

Serves global market

Offices12 records

Each record includes

City, Country, Type, Description, Source

Keyword5 values
AI cloud infrastructure, GPU compute services, generative AI inference, open-source AI platform, AI model deployment
Industry5 codes
1AI Compute Cloud & GPU-as-a-Service
CodeHDAAAAAKPrimaryYes
2Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem)
CodeHDAEANACPrimaryNo
3AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers)
CodeHDAAAAAIPrimaryNo
4End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management)
CodeHDAEANAAPrimaryNo
5AI Integration & Orchestration Platforms (Connectors, Workflow, iPaaS for AI)
CodeHDAEANAIPrimaryNo
NAICS code4 codes
  • Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services5182
  • Software Publishers5132
  • Computer Systems Design and Related Services54151
  • Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services518
SIC code3 codes
  • Services-Computer Programming, Data Processing, Etc.7370
  • Services-Prepackaged Software7372
  • Services-Computer Integrated Systems Design7373
Product category
AI Cloud Infrastructure
Social media profiles2 records
GTM motion2 records

Each record includes

Type, Description, Source

Revenue model7 records
1GPU Cluster Compute (On-Demand & Reserved)
TypeUsage Based
Description

Hourly billing for GPU compute capacity on NVIDIA H100, H200, B200, GB200 NVL72, and GB300 NVL72 clusters. On-demand pricing for flexible usage; reserved capacity at lower hourly rates with upfront commitment for 1-6 month terms. Scales from single-node 8-GPU to 4,000+ GPU deployments. Used for model training, fine-tuning, and distributed AI research workloads.

theaiinsider.tech
2Serverless Inference API (Per-Token)
TypeUsage Based
Description

Pay-per-token pricing for real-time inference APIs across 200+ open-source models including DeepSeek V4 Pro, Llama 4, Qwen3, MiniMax M3, Kimi K2, NVIDIA Nemotron 3, GLM-5, and more. Customers pay based on input and output token volume with no infrastructure management required. Up to 50% cost savings vs. real-time API pricing for batch workloads.

theaiinsider.tech
3Batch Inference (Async)
TypeUsage Based
Description

Asynchronous batch processing of up to 30 billion enqueued tokens per model per user at up to 50% discount vs. real-time API pricing. Customers upload JSONL files and receive processed results without real-time latency requirements.

theaiinsider.tech
4Dedicated Model Inference
TypeSubscription Recurring
Description

Reserved, isolated compute resources for dedicated inference endpoints. Customers deploy specific models on dedicated infrastructure with SLA-backed performance guarantees, providing predictable pricing and consistent latency for production workloads.

theaiinsider.tech
5Dedicated Container Inference
TypeManaged Services
Description

GPU infrastructure for running custom containers and generative media models (video, audio, image) with the Sprocket SDK. Supports multi-GPU orchestration and elastic autoscaling. Revenue from container deployment and per-request pricing with no minimum commitments.

theaiinsider.tech
6Fine-Tuning Services
TypeManaged Services
Description

Fine-tuning of open-source models using customer data. Supports LoRA and full fine-tuning across model sizes including 100B+ parameter models. Pricing based on training compute time with cost estimation upfront.

theaiinsider.tech
7AI Factory (Custom Infrastructure)
TypeManaged Services
Description

Custom frontier-scale AI infrastructure deployments for enterprise customers needing 1,000-100,000+ GPU capacity. Reserved long-term capacity with custom configuration for trillion-parameter model training and large-scale inference operations.

theaiinsider.tech
Marketing channels9 records

Each record includes

Title, Type, Stage, Description, Source

Distribution channels5 records

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components6 values
Infrastructure, Technology or R&D, Personnel, Marketing or Sales, Operations, Supply Chain
Pricing details14 tiers
1H100 GPU Cluster — On-Demand
ModelUsage-basedBilling cadencePay-as-you-go
Notes

$4.79/hr per GPU. Scale: 8 to 256 GPUs. NVIDIA HGX H100 SXM (80GB).

together.ai
2H100 GPU Cluster — Reserved
ModelUsage-basedBilling cadenceMulti-year contract
Notes

Starting at $3.29/hr per GPU. Up to 6-month commitment, pay upfront. Scale: 8 to 256 GPUs.

together.ai
3H200 GPU Cluster — On-Demand
ModelUsage-basedBilling cadencePay-as-you-go
Notes

$5.99/hr per GPU. NVIDIA HGX H200 (141GB). Scale: 256 to 1,000 GPUs.

together.ai
4H200 GPU Cluster — Reserved
ModelUsage-basedBilling cadenceMulti-year contract
Notes

Starting at $3.99/hr per GPU. Up to 6-month commitment.

together.ai
5B200 GPU Cluster — On-Demand
ModelUsage-basedBilling cadencePay-as-you-go
Notes

$8.19/hr per GPU. NVIDIA HGX B200. Scale: 256 to 1,000+ GPUs.

together.ai
6B200 GPU Cluster — Reserved
ModelUsage-basedBilling cadenceMulti-year contract
Notes

Starting at $6.79/hr per GPU. Up to 6-month commitment.

together.ai
7GB200 NVL72 / GB300 NVL72 — Custom
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Contact sales for pricing. Reserved capacity available. Scale: 512 to 1,000+ GPUs.

together.ai
8DeepSeek V4 Pro — Serverless Inference
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Cached: $0.20/M tokens; Input: $1.74/M tokens; Output: $3.48/M tokens. Function calling, JSON mode, reasoning modes.

together.ai
9MiniMax M3 — Serverless Inference
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Cached: $0.06/M tokens; Input: $0.30/M tokens; Output: $1.20/M tokens. Vision, reasoning, function calling.

together.ai
10NVIDIA Nemotron 3 Ultra — Serverless Inference
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Cached: $0.60/M tokens; Input: $3.60/M tokens; Output: pricing varies. Function calling, reasoning, code.

together.ai
11GPT-oss-120B — Serverless Inference
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Input: $0.15/M tokens. 120B parameters. JSON mode, reasoning.

together.ai
12Batch Inference — 50% Off Many Top Models
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Up to 50% cost savings vs. real-time API. Scale to 30B enqueued tokens per model. <24h processing SLA. Models include DeepSeek, Llama, Qwen, Kimi.

together.ai
13Instant Clusters — Self-Service GPU Provisioning
ModelUsage-basedBilling cadencePay-as-you-go
Notes

$1.76 to $5.50/hr depending on hardware configuration and commitment level. Supports Terraform and SkyPilot. Preloaded drivers and networking.

siliconangle.com
14Batch Inference — Async Processing
ModelUsage-basedBilling cadencePay-as-you-go
Notes

Up to 50% savings vs real-time API. 30B tokens per model enqueued. <24h SLA (often within hours). No minimum volume required.

together.ai
GTM typeB2B
B2B
Offering typeSoftware
Software
Core offering1 text field

Together AI operates an AI Native Cloud platform that provides GPU-accelerated compute, inference APIs, fine-tuning, and proprietary systems research for serving open-source generative AI models at scale. Customers access 200+ open-source models from 40+ providers through serverless, batch, dedicated, and container deployment modes, backed by research-driven kernel optimizations that deliver higher throughput and lower cost than closed-model alternatives.

Differentiator
Functional benefit
Problem solved
Quantifiable outcome1 of 12 values shown
  • Up to 60% lower inference costs vs. proprietary AI APIs, with 6x to 60x savings reported by enterprise customers switching from closed models
+11 more records
Product overview1 text field

Together AI is an AI Native Cloud platform offering a comprehensive full-stack portfolio of inference, compute, and model shaping products. The core inference layer includes Serverless Inference (auto-scaling pay-per-token API), Batch Inference (asynchronous processing up to 30B tokens), Dedicated Model Inference (reserved GPU infrastructure with proprietary inference engine), and Dedicated Container Inference (for generative media). The accelerated compute layer provides GPU Clusters (self-serve H100/H200/B200/GB200/GB300 with InfiniBand) and Frontier AI Factory (custom 1000+ GPU deployments for trillion-parameter models). Supporting products include Sandbox (fast code environments), Managed Storage (optimized object/parallel storage), and Fine-Tuning (LoRA and full fine-tuning for 100B+ models). Research products including FlashAttention-4, ATLAS, ThunderKittens, Megakernel, and Together Kernel Collection power the platform with memory-efficient attention, adaptive speculative decoding, and custom CUDA kernels achieving up to 4x inference speedups and 90% training acceleration. The platform serves thousands of customers including Cursor, Decagon, Cognition, and The Washington Post.

Product and service9 records
1Serverless Inference
CategoryInference API
Description

Fully managed pay-per-token inference API with automatic scaling for 200+ open-source models, providing high-performance inference without infrastructure management or long-term commitments. Powers real-time AI workloads for AI-native startups, developers, and enterprise customers.

2Batch Inference
CategoryBatch Processing API
Description

Asynchronous batch processing service that scales to 30 billion enqueued tokens per model per user with up to 50% cost savings versus real-time API. Jobs finish within a 24-hour SLA, often within hours, for massive inference workloads.

3Dedicated Model Inference
CategoryDedicated Inference
Description

Reserved, isolated compute resources running Together AI's proprietary inference engine (ATLAS adaptive speculative decoding) for production workloads needing consistent latency, control, and best economics with SLA-backed performance guarantees.

4Dedicated Container Inference
CategoryContainer Inference
Description

GPU infrastructure purpose-built for generative media workloads (video, audio, image) running on the Sprocket SDK with multi-GPU orchestration and elastic autoscaling for 10x traffic surges. Multi-cluster scaling handles viral demand for video, audio, and avatar generation models.

5GPU Clusters
CategoryAccelerated Compute
Description

Self-serve GPU clusters at scale with bare-metal performance, InfiniBand networking, and managed orchestration supporting NVIDIA H100, H200, B200, GB200, and GB300 GPUs. Flexible on-demand pricing ($1.76-$8.19/hr per GPU) and reserved capacity for 1-6 month commitments, scaling from 8 to 4,000+ GPUs.

6Frontier AI Factory
CategoryCustom AI Infrastructure
Description

Custom infrastructure at frontier scale for trillion-parameter model training and large-scale inference operations, configured for 1,000-100,000+ GPU capacity with NVIDIA Blackwell GPUs, managed orchestration, self-healing infrastructure, and expert support. Reserved long-term capacity for customers including Hedra.

7Fine-Tuning
CategoryModel Customization
Description

Fine-tuning service for customizing open-source models with user data, supporting LoRA and full fine-tuning for 100B+ parameter models. Includes vision fine-tuning, tool-calling training, and multi-GPU distributed training with upfront cost estimation.

8Sandbox
CategoryDeveloper Environment
Description

Fast, secure code sandboxes at scale for AI development environments featuring 2.7-second cold starts, 500ms snapshot resumes, VM cloning, and DevContainer support. Powers AI coding workflows including HeroUI's UI component library with 98% lower preview cold starts.

9Managed Storage
CategoryStorage
Description

High-performance managed storage for AI-native workloads providing object storage and parallel filesystems optimized to keep GPUs fed during training and inference. Includes zero egress fees for data movement between regions.

Scale indicator21 records

Each record includes

Type, Value, Description, Source

Partnership7 partners
Strategic tierCoreTypeTechnology or IntegrationAnnounced on2026-07-01
Description

Pegatron established a strategic collaboration with Together AI and 5C to deliver large-scale AI infrastructure using NVIDIA GB300 NVL72 and HGX B200 liquid-cooled rack deployments in US data centers (Texas and Maryland). Pegatron provides manufacturing and deployment expertise for NVIDIA-powered AI factory infrastructure.

Strategic tierCoreTypeTechnology or IntegrationAnnounced on2026-06-05
Description

Together AI listed among NVIDIA Agent Toolkit partners at GTC 2026, alongside CoreWeave, Fireworks, and cloud infrastructure providers integrating NVIDIA's open-source agent development platform.

Strategic tierCoreTypeTechnology or IntegrationAnnounced on2026-06-04
Description

Rumble (rebranded as RUM Group Inc.) signed a $270M multi-year GPU cloud agreement in June 2026 to provide dedicated NVIDIA HGX Blackwell B300 GPU capacity to Together AI. Rumble acquired Northern Data's GPU estate (~22,000 NVIDIA H100/H200 GPUs) to fulfill the agreement. This expands Together AI's global GPU footprint with Blackwell-class capacity for large-scale AI training, fine-tuning, and inference workloads.

Strategic tierStrategicTypeStrategic or Co-development PartnerAnnounced on2026-04-30
Description

Together AI partnered with Adaption (co-founded by former Cohere and Google DeepMind leaders Sara Hooker and Sudip Roy) to integrate Together Fine-Tuning capabilities natively into Adaption's data management platform. Enables seamless transition from data optimization to model fine-tuning. Adaption reports 82% average increase in data quality for early users.

Strategic tierStrategicTypeStrategic or Co-development PartnerAnnounced on2026-04-27
Description

Together AI joined the U.S. DOE's Genesis Mission, a national initiative uniting 17 National Laboratories, academia, and private industry to build an integrated AI discovery platform aimed at doubling American scientific productivity within a decade. Together AI contributes FlashAttention and high-performance inference infrastructure to support frontier research in energy and national security.

Strategic tierMinorTypeStrategic or Co-development PartnerAnnounced on2026-03-05
Description

Collinear partnered with Together AI to integrate TraitBasis (method for generating realistic simulated users) into the Together Evals platform, enabling builders to test AI models against realistic user behaviors.

Strategic tierCoreTypeTechnology or IntegrationAnnounced on2026-01-01
Description

Together AI integrates with Hugging Face's model hosting ecosystem. Customers can deploy virtually any model from Hugging Face with minimal friction via Dedicated Container Inference (DCI) offering. Together AI listed among top open-source AI model API providers alongside Hugging Face's open-source ecosystem.

Recent move8 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight7 records

Each record includes

Type, Description

Peers10 records
TypeDirect peer
Description

CoreWeave is a specialized GPU cloud provider offering NVIDIA-powered compute, inference, and AI factory services to AI labs and enterprises. It is the closest direct competitor to Together AI's GPU Clusters and Dedicated Inference offerings and competes head-to-head on H100/H200/B200 capacity.

TypeDirect peer
Description

Lambda operates GPU cloud instances (H100, H200, B200 clusters) and dedicated AI training/inference infrastructure, directly overlapping Together AI's GPU Clusters and AI Factory products for AI-native startups and enterprise research teams.

TypeDirect peer
Description

Fireworks AI provides serverless and dedicated inference APIs for open-source and fine-tuned models with proprietary optimization techniques. It is a direct competitor in the open-model inference API space, targeting the same AI-native developer and enterprise segments as Together AI.

TypeDirect peer
Description

Replicate runs a cloud platform for running and deploying open-source AI models via API, with a focus on generative media (image, video, audio). It overlaps with Together AI's Serverless Inference and Dedicated Container Inference offerings for the same open-model developer ecosystem.

TypeDirect peer
Description

Anyscale, creator of Ray, offers AI compute and inference platform services on GPU clusters and is repositioning around production AI workloads. It competes with Together AI on enterprise inference, fine-tuning, and large-scale training infrastructure.

TypeEmerging player
Description

Modal provides serverless GPU compute for AI inference and batch workloads, with strong developer ergonomics. It targets a similar developer audience as Together AI's Serverless Inference and Instant Clusters products, though with a smaller model catalog and footprint.

TypeDirect peer
Description

Hugging Face hosts open-source models and offers Inference API, Endpoints, and Spaces for deploying AI models. It is both a partner and a competitor to Together AI, particularly around open-model hosting and developer-facing inference APIs.

TypeEmerging player
Description

Nebius is rebuilding an AI-focused GPU cloud infrastructure business out of Yandex's hardware estate, offering NVIDIA GPU clusters and managed AI services. It overlaps with Together AI's GPU Clusters and AI Factory offerings, particularly in Europe.

TypeEmerging player
Description

Crusoe builds large-scale, energy-optimized data centers for AI compute, including NVIDIA GPU clusters and managed AI cloud services. It competes with Together AI in GPU cluster capacity and AI Factory-style deployments for large training runs.

TypeBroad incumbent
Description

AWS offers EC2 GPU instances (P5/H100, P6/B200), Bedrock managed inference, and SageMaker AI platforms. As a broad incumbent hyperscaler, AWS competes with Together AI on raw GPU capacity and increasingly on managed inference for both AI-native and enterprise customers.

Market position
Strengths5 records

Each record includes

Headline, Details, Source

Weaknesses5 records

Each record includes

Headline, Details, Source

Competitive moat6 records

Each record includes

Type, Details

Key risks6 records

Each record includes

Headline, Details, Source

Key highlights7 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Named customers20 records

Each record includes

Name, Industry, Type, Use case, Source, UUID

Segment4 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile3 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
Yes
API detail
Has APIbool
Yes

Docs URL, Description

Integration3 records

Each record includes

Title, Type, Description, Source

AI capability12 records

Each record includes

Type, Description, Source

AI maturity
App detail

Has app

Feature9 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
Profiles8 records

Each record includes

Name, Designation, Designation category, Overview, Profile commentary, Source

Subsidiaries1 record

Each record includes

Name, Acquired on, Relationship type, Type, Business focus

Compliance2 records

Each record includes

Name, Class, Description

Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds5 records

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors41 records

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A2 records

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Together AI

AI Cloud Infrastructuretogether.ai

Together AI operates an AI Native Cloud platform delivering GPU-accelerated compute, serverless and dedicated inference, fine-tuning, and proprietary kernel-level optimizations for 200+ open-source AI models. It serves AI-native startups, generative media companies, and enterprises across 25+ global cities.

What Together AI does

Together AI is an AI Native Cloud platform founded in 2022 that provides GPU-accelerated compute infrastructure, inference APIs, fine-tuning, and research-driven optimizations for open-source AI models. The platform supports NVIDIA H100, H200, B200, GB200 NVL72, and GB300 NVL72 GPU clusters with InfiniBand networking, and hosts more than 200 open-source models from providers including DeepSeek, Meta (Llama), Qwen, Google, OpenAI, Mistral, NVIDIA, Moonshot, and MiniMax. The technical differentiator is the Together Kernel Collection—custom CUDA kernels co-developed with Chief Scientist Tri Dao (creator of FlashAttention)—together with ATLAS adaptive speculative decoding, ThunderKittens (an embedded DSL for GPU kernels), Megakernel (single-kernel model forward passes), and FlashAttention-4, which together deliver up to 90% faster training and 2x faster inference versus standard stacks.

The revenue model is built around four usage-based and managed-services streams: per-token pricing on serverless and batch inference APIs (with batch offering up to 50% discount versus real-time), hourly on-demand and reserved GPU cluster compute, dedicated inference deployments with SLA-backed performance, and managed fine-tuning for 100B+ parameter models. Distribution combines API-first product-led growth with a 1M+ developer base and direct enterprise sales for large accounts, with ISO 27001:2022, SOC 2 Type II, and HIPAA-aligned storage available for regulated buyers. The company serves a horizontal customer mix including AI-native startups (Cursor, Cognition, Decagon), generative media platforms (Pika, Hedra, Krea), and enterprises (Salesforce, Zoom, The Washington Post, LG AI Research), reporting 6x to 60x cost savings for customers switching from closed-model APIs.

The company is headquartered in San Francisco with infrastructure deployed across more than 25 cities in the U.S., Europe, and Asia/Middle East, and has raised over $1.3 billion across five funding rounds since 2023. A Series C of $800M led by Aramco Ventures in July 2026 set a post-money valuation of $8.3 billion, with the company reporting annual bookings exceeding $1.15 billion and an annualized revenue run rate of approximately $1 billion as of February 2026. Strategic positioning as an NVIDIA Preferred Partner and participant in the U.S. DOE Genesis Mission underpins both commercial and government-sector reach.

Together AI firmographics

Firmographics
Name
Together AI
Legal name
Together AI
Website
https://together.ai
Company type
Private
Founded year
2022
Operating status
Operating
Headcount range
251–500 employees
Short description
Together AI operates an AI Native Cloud platform delivering GPU-accelerated compute, serverless and dedicated inference, fine-tuning, and proprietary kernel-level optimizations for 200+ open-source AI models. It serves AI-native startups, generative media companies, and enterprises across 25+ global cities.
Ownership category
akta.pro rank

Together AI industry classification

Industry
Product category
AI Cloud Infrastructure
NAICS
Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182), Software Publishers (5132), Computer Systems Design and Related Services (54151), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (518)
SIC
Services-Computer Programming, Data Processing, Etc. (7370), Services-Prepackaged Software (7372), Services-Computer Integrated Systems Design (7373)
akta.pro primary industry
AI Compute Cloud & GPU-as-a-Service (HDAAAAAK)
akta.pro secondary industries
Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC), AI Compiler, Runtime & Kernel Optimization Software (CUDA/ROCm/XLA, graph compilers) (HDAAAAAI), End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management) (HDAEANAA), AI Integration & Orchestration Platforms (Connectors, Workflow, iPaaS for AI) (HDAEANAI)

Keywords

  • AI cloud infrastructure
  • GPU compute services
  • Generative AI inference
  • Open-source AI platform
  • AI model deployment

Where Together AI is headquartered

Location

Headquarters

HQ city
San Francisco
HQ country
United States
HQ region
North America

Offices12 records

Markets served

Together AI business model

Business model
GTM type
B2B
Offering type
Software
Cost components
Infrastructure, Technology or R&D, Personnel, Marketing or Sales, Operations, Supply Chain

Revenue model

  1. GPU Cluster Compute (On-Demand & Reserved): Hourly billing for GPU compute capacity on NVIDIA H100, H200, B200, GB200 NVL72, and GB300 NVL72 clusters. On-demand pricing for flexible usage; reserved capacity at lower hourly rates with upfront commitment for 1-6 month terms. Scales from single-node 8-GPU to 4,000+ GPU deployments. Used for model training, fine-tuning, and distributed AI research workloads.
  2. Serverless Inference API (Per-Token): Pay-per-token pricing for real-time inference APIs across 200+ open-source models including DeepSeek V4 Pro, Llama 4, Qwen3, MiniMax M3, Kimi K2, NVIDIA Nemotron 3, GLM-5, and more. Customers pay based on input and output token volume with no infrastructure management required. Up to 50% cost savings vs. real-time API pricing for batch workloads.
  3. Batch Inference (Async): Asynchronous batch processing of up to 30 billion enqueued tokens per model per user at up to 50% discount vs. real-time API pricing. Customers upload JSONL files and receive processed results without real-time latency requirements.
  4. Dedicated Model Inference: Reserved, isolated compute resources for dedicated inference endpoints. Customers deploy specific models on dedicated infrastructure with SLA-backed performance guarantees, providing predictable pricing and consistent latency for production workloads.
  5. Dedicated Container Inference: GPU infrastructure for running custom containers and generative media models (video, audio, image) with the Sprocket SDK. Supports multi-GPU orchestration and elastic autoscaling. Revenue from container deployment and per-request pricing with no minimum commitments.
  6. Fine-Tuning Services: Fine-tuning of open-source models using customer data. Supports LoRA and full fine-tuning across model sizes including 100B+ parameter models. Pricing based on training compute time with cost estimation upfront.
  7. AI Factory (Custom Infrastructure): Custom frontier-scale AI infrastructure deployments for enterprise customers needing 1,000-100,000+ GPU capacity. Reserved long-term capacity with custom configuration for trillion-parameter model training and large-scale inference operations.

Pricing tiers

ModelBillingPrice
Usage-basedPay-as-you-goH100 GPU Cluster — On-Demand
Usage-basedMulti-year contractH100 GPU Cluster — Reserved
Usage-basedPay-as-you-goH200 GPU Cluster — On-Demand
Usage-basedMulti-year contractH200 GPU Cluster — Reserved
Usage-basedPay-as-you-goB200 GPU Cluster — On-Demand
Usage-basedMulti-year contractB200 GPU Cluster — Reserved
Usage-basedPay-as-you-goGB200 NVL72 / GB300 NVL72 — Custom
Usage-basedPay-as-you-goDeepSeek V4 Pro — Serverless Inference
Usage-basedPay-as-you-goMiniMax M3 — Serverless Inference
Usage-basedPay-as-you-goNVIDIA Nemotron 3 Ultra — Serverless Inference
Usage-basedPay-as-you-goGPT-oss-120B — Serverless Inference
Usage-basedPay-as-you-goBatch Inference — 50% Off Many Top Models
Usage-basedPay-as-you-goInstant Clusters — Self-Service GPU Provisioning
Usage-basedPay-as-you-goBatch Inference — Async Processing

Go-to-market motion2 records

Distribution channels5 records

Marketing channels9 records

Together AI product offering

Product offering

Core offering

Together AI operates an AI Native Cloud platform that provides GPU-accelerated compute, inference APIs, fine-tuning, and proprietary systems research for serving open-source generative AI models at scale. Customers access 200+ open-source models from 40+ providers through serverless, batch, dedicated, and container deployment modes, backed by research-driven kernel optimizations that deliver higher throughput and lower cost than closed-model alternatives.

Product overview

Together AI is an AI Native Cloud platform offering a comprehensive full-stack portfolio of inference, compute, and model shaping products. The core inference layer includes Serverless Inference (auto-scaling pay-per-token API), Batch Inference (asynchronous processing up to 30B tokens), Dedicated Model Inference (reserved GPU infrastructure with proprietary inference engine), and Dedicated Container Inference (for generative media). The accelerated compute layer provides GPU Clusters (self-serve H100/H200/B200/GB200/GB300 with InfiniBand) and Frontier AI Factory (custom 1000+ GPU deployments for trillion-parameter models). Supporting products include Sandbox (fast code environments), Managed Storage (optimized object/parallel storage), and Fine-Tuning (LoRA and full fine-tuning for 100B+ models). Research products including FlashAttention-4, ATLAS, ThunderKittens, Megakernel, and Together Kernel Collection power the platform with memory-efficient attention, adaptive speculative decoding, and custom CUDA kernels achieving up to 4x inference speedups and 90% training acceleration. The platform serves thousands of customers including Cursor, Decagon, Cognition, and The Washington Post.

Differentiator

Problem solved

Functional benefit

Products and services

  • Serverless Inference Fully managed pay-per-token inference API with automatic scaling for 200+ open-source models, providing high-performance inference without infrastructure management or long-term commitments. Powers real-time AI workloads for AI-native startups, developers, and enterprise customers.
  • Batch Inference Asynchronous batch processing service that scales to 30 billion enqueued tokens per model per user with up to 50% cost savings versus real-time API. Jobs finish within a 24-hour SLA, often within hours, for massive inference workloads.
  • Dedicated Model Inference Reserved, isolated compute resources running Together AI's proprietary inference engine (ATLAS adaptive speculative decoding) for production workloads needing consistent latency, control, and best economics with SLA-backed performance guarantees.
  • Dedicated Container Inference GPU infrastructure purpose-built for generative media workloads (video, audio, image) running on the Sprocket SDK with multi-GPU orchestration and elastic autoscaling for 10x traffic surges. Multi-cluster scaling handles viral demand for video, audio, and avatar generation models.
  • GPU Clusters Self-serve GPU clusters at scale with bare-metal performance, InfiniBand networking, and managed orchestration supporting NVIDIA H100, H200, B200, GB200, and GB300 GPUs. Flexible on-demand pricing ($1.76-$8.19/hr per GPU) and reserved capacity for 1-6 month commitments, scaling from 8 to 4,000+ GPUs.
  • Frontier AI Factory Custom infrastructure at frontier scale for trillion-parameter model training and large-scale inference operations, configured for 1,000-100,000+ GPU capacity with NVIDIA Blackwell GPUs, managed orchestration, self-healing infrastructure, and expert support. Reserved long-term capacity for customers including Hedra.
  • Fine-Tuning Fine-tuning service for customizing open-source models with user data, supporting LoRA and full fine-tuning for 100B+ parameter models. Includes vision fine-tuning, tool-calling training, and multi-GPU distributed training with upfront cost estimation.
  • Sandbox Fast, secure code sandboxes at scale for AI development environments featuring 2.7-second cold starts, 500ms snapshot resumes, VM cloning, and DevContainer support. Powers AI coding workflows including HeroUI's UI component library with 98% lower preview cold starts.
  • Managed Storage High-performance managed storage for AI-native workloads providing object storage and parallel filesystems optimized to keep GPUs fed during training and inference. Includes zero egress fees for data movement between regions.

Quantifiable outcome

  • Up to 60% lower inference costs vs. proprietary AI APIs, with 6x to 60x savings reported by enterprise customers switching from closed models
  • +11 more outcomes

Companies that use Together AI

Customer profile

Named customers20 records

Segments4 records

Ideal customer profiles3 records

Together AI technology and API

Technology

Technology focussed Yes

API detail

Has API
Yes
API docs
API detail

Core technology

AI maturity

App detail

Integration3 records

AI capability12 records

Feature9 records

Together AI partnerships and signals

Strategic signal

Partnerships

Seven partnerships are on record, tiered core, strategic and minor.

  • PegatroncoreTechnology or Integration · 1 July 2026Pegatron established a strategic collaboration with Together AI and 5C to deliver large-scale AI infrastructure using NVIDIA GB300 NVL72 and HGX B200 liquid-cooled rack deployments in US data centers (Texas and Maryland). Pegatron provides manufacturing and deployment expertise for NVIDIA-powered AI factory infrastructure.
  • NVIDIA (Agent Toolkit partner)coreTechnology or Integration · 5 June 2026Together AI listed among NVIDIA Agent Toolkit partners at GTC 2026, alongside CoreWeave, Fireworks, and cloud infrastructure providers integrating NVIDIA's open-source agent development platform.
  • Rumble (Quake AI)coreTechnology or Integration · 4 June 2026Rumble (rebranded as RUM Group Inc.) signed a $270M multi-year GPU cloud agreement in June 2026 to provide dedicated NVIDIA HGX Blackwell B300 GPU capacity to Together AI. Rumble acquired Northern Data's GPU estate (~22,000 NVIDIA H100/H200 GPUs) to fulfill the agreement. This expands Together AI's global GPU footprint with Blackwell-class capacity for large-scale AI training, fine-tuning, and inference workloads.
  • AdaptionstrategicStrategic or Co-development Partner · 30 April 2026Together AI partnered with Adaption (co-founded by former Cohere and Google DeepMind leaders Sara Hooker and Sudip Roy) to integrate Together Fine-Tuning capabilities natively into Adaption's data management platform. Enables seamless transition from data optimization to model fine-tuning. Adaption reports 82% average increase in data quality for early users.
  • U.S. Department of EnergystrategicStrategic or Co-development Partner · 27 April 2026Together AI joined the U.S. DOE's Genesis Mission, a national initiative uniting 17 National Laboratories, academia, and private industry to build an integrated AI discovery platform aimed at doubling American scientific productivity within a decade. Together AI contributes FlashAttention and high-performance inference infrastructure to support frontier research in energy and national security.
  • CollinearminorStrategic or Co-development Partner · 5 March 2026Collinear partnered with Together AI to integrate TraitBasis (method for generating realistic simulated users) into the Together Evals platform, enabling builders to test AI models against realistic user behaviors.
  • Hugging FacecoreTechnology or Integration · 1 January 2026Together AI integrates with Hugging Face's model hosting ecosystem. Customers can deploy virtually any model from Hugging Face with minimal friction via Dedicated Container Inference (DCI) offering. Together AI listed among top open-source AI model API providers alongside Hugging Face's open-source ecosystem.

Scale indicators21 records

Recent moves8 records

Expansion highlights7 records

Together AI competitors and assessment

Company assessment

Direct peers

  • CoreWeave: CoreWeave is a specialized GPU cloud provider offering NVIDIA-powered compute, inference, and AI factory services to AI labs and enterprises. It is the closest direct competitor to Together AI's GPU Clusters and Dedicated Inference offerings and competes head-to-head on H100/H200/B200 capacity.
  • Lambda Labs: Lambda operates GPU cloud instances (H100, H200, B200 clusters) and dedicated AI training/inference infrastructure, directly overlapping Together AI's GPU Clusters and AI Factory products for AI-native startups and enterprise research teams.
  • Fireworks AI: Fireworks AI provides serverless and dedicated inference APIs for open-source and fine-tuned models with proprietary optimization techniques. It is a direct competitor in the open-model inference API space, targeting the same AI-native developer and enterprise segments as Together AI.
  • Replicate: Replicate runs a cloud platform for running and deploying open-source AI models via API, with a focus on generative media (image, video, audio). It overlaps with Together AI's Serverless Inference and Dedicated Container Inference offerings for the same open-model developer ecosystem.
  • Anyscale: Anyscale, creator of Ray, offers AI compute and inference platform services on GPU clusters and is repositioning around production AI workloads. It competes with Together AI on enterprise inference, fine-tuning, and large-scale training infrastructure.
  • Hugging Face: Hugging Face hosts open-source models and offers Inference API, Endpoints, and Spaces for deploying AI models. It is both a partner and a competitor to Together AI, particularly around open-model hosting and developer-facing inference APIs.

Emerging players

  • Modal: Modal provides serverless GPU compute for AI inference and batch workloads, with strong developer ergonomics. It targets a similar developer audience as Together AI's Serverless Inference and Instant Clusters products, though with a smaller model catalog and footprint.
  • Nebius: Nebius is rebuilding an AI-focused GPU cloud infrastructure business out of Yandex's hardware estate, offering NVIDIA GPU clusters and managed AI services. It overlaps with Together AI's GPU Clusters and AI Factory offerings, particularly in Europe.
  • Crusoe: Crusoe builds large-scale, energy-optimized data centers for AI compute, including NVIDIA GPU clusters and managed AI cloud services. It competes with Together AI in GPU cluster capacity and AI Factory-style deployments for large training runs.

Broad incumbents

  • Amazon Web Services (AWS): AWS offers EC2 GPU instances (P5/H100, P6/B200), Bedrock managed inference, and SageMaker AI platforms. As a broad incumbent hyperscaler, AWS competes with Together AI on raw GPU capacity and increasingly on managed inference for both AI-native and enterprise customers.

Market position

Strengths5 records

Weaknesses5 records

Competitive moat6 records

Key risks6 records

Key highlights7 records

Customer concentration

Together AI social profiles

Digital presence

Together AI compliance and trust

Trust signal

Compliance2 records

Together AI financial estimates

Financial estimate

Revenue estimate

Valuation estimate

Together AI leadership team

Management profile

Number of profiles

Profiles8 records

Together AI subsidiaries and ownership

Company hierarchy

Subsidiaries1 record

Together AI funding detail

Funding detail

Funding overview

Funding rounds5 records

Investors41 records

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

Together AI M&A and investment

M&A and investment

M&A2 records

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about Together AI

What does Together AI do?

Together AI operates an AI Native Cloud platform that provides GPU-accelerated compute, inference APIs, fine-tuning, and proprietary systems research for serving open-source generative AI models at scale. Customers access 200+ open-source models from 40+ providers through serverless, batch, dedicated, and container deployment modes, backed by research-driven kernel optimizations that deliver higher throughput and lower cost than closed-model alternatives.

Is Together AI a public or private company?

Together AI is a private company. It is classified as venture growth investor backed and is currently operating.

When was Together AI founded?

Together AI was founded in 2022. It employs 251 to 500 people.

Where is Together AI based?

Together AI is headquartered in San Francisco, United States, in the North America region.

How does Together AI make money?

Seven revenue lines are on record. GPU Cluster Compute (On-Demand & Reserved) is the primary driver. The others are serverless Inference API (Per-Token), batch Inference (Async), dedicated Model Inference, dedicated Container Inference, fine-Tuning Services and AI Factory (Custom Infrastructure).

Who are Together AI's main competitors?

Direct peers on record are CoreWeave, Lambda Labs, Fireworks AI, Replicate, Anyscale and Hugging Face. Emerging players are Modal, Nebius and Crusoe. Amazon Web Services (AWS) is listed as a broad incumbent.

Does Together AI have an API?

Yes. Together AI offers a public REST API for inference, fine-tuning, GPU clusters, and model deployment. The API supports serverless inference, batch inference, dedicated model inference, fine-tuning jobs, and cluster management. Authentication via API keys, with endpoints available at api.together.ai. Documentation available at docs.together.ai with SDK support in Python. Developer documentation is at docs.together.ai.

What industry is Together AI in?

Together AI's product category is AI Cloud Infrastructure. Its primary akta.pro industry code is HDAAAAAK, AI Compute Cloud & GPU-as-a-Service, with a secondary code of HDAEANAC, Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem). Its NAICS code is 5182 and its SIC code is 7370.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales
Live signals
HackerNoonOpen Weight AI Is Moving the Competitive Advantage From Models to InferenceMeta released Muse Glimmer, a compact open-weight model for single-GPU devices, and plans to open Muse Spark 1.2 weights. IBM and Together AI signed a $240 million agreement to build an inference cluster on IBM Cloud. The shift moves competitive advantage from model training to inference deployment and operation.MarkTechPostMeet Together Link: A Free CLI That Runs Open Models Like Kimi K3 and GLM 5.3 Inside Claude Code, Codex, and OpenCodeTogether AI released Together Link, a free MIT-licensed CLI in beta that connects coding agents to open models on Together AI. It installs with one command, routes tasks between models like Kimi K3 and GLM 5.3, and claims over 50% savings versus all-Opus 5.5 sessions. The tool is available on macOS and Linux during beta.Pulse 2.0Clockwork.io Raises $31 Million To Expand AI Infrastructure Fault-Tolerance PlatformClockwork.io raised $31 million in new funding, bringing total funding to $73 million, to expand its fault-tolerance software for AI infrastructure. The company plans to accelerate deployment across AI training, inference, and reinforcement learning, and announced production deployments at LinkedIn and Together AI. It also introduced new TorchPass capabilities for multi-node snapshots and asynchronous checkpoints.PR NewswireClockwork.io recauda 31 millones de dólaresClockwork.io raised $31 million in a new funding round, with production deployments at LinkedIn, Together AI, and WhiteFiber. The company also announced new TorchPass features that capture AI workload state without code changes. Total funding now reaches $73 million.TechCrunchOpen or closed AI? Learn what to build on at Disrupt 2026TechCrunch Disrupt 2026 will feature four sessions on AI model choices, covering multi-model applications, ownership, and hardware. Speakers include leaders from Together AI, Pathway, Oumi, and Nvidia, discussing trade-offs between open and proprietary models. The event runs October 13-15 in San Francisco.Third NewsClockwork.io Secures $31M Funding, Empowering AI With Unmatched Resilience TechnologyClockwork.io raised $31 million in funding to enhance AI workload fault-tolerance, with cumulative funding now at $73 million. Its products LinkPass and TorchPass reroute traffic around failures and enable rapid recovery, adopted by LinkedIn, Together AI, and WhiteFiber. The funding will support development for AI training, inference, and reinforcement learning.BitlyClockwork.io Raises $31M as LinkedIn, Together AI and WhiteFiber Adopt Its Resilience Software to Stop Wasting GPU-HoursClockwork.io raised $31 million in funding, with production deployments at LinkedIn and Together AI and expanded adoption by WhiteFiber. The company announced new TorchPass capabilities, including platform snapshots and fast application checkpoints, to preserve AI workload progress without code changes. The round brings total funding to $73 million.Stock TitanClockwork.io Raises $31M for AI Fault-Tolerance SoftwareClockwork.io raised $31 million in funding, with production deployments at LinkedIn and Together AI and expanded adoption by WhiteFiber. The company introduced new TorchPass capabilities, including platform snapshots and fast asynchronous checkpoints, to preserve AI workload progress. It will use the capital to expand enterprise adoption and scale delivery through cloud partners.Tech Funding NewsClockwork.io lands $31M to eliminate wasted GPU-hours with AI infrastructure fault tolerance — TFNClockwork.io raised $31 million in new funding, bringing total capital to $73 million, to support its fault-tolerance platform for AI clusters. The company's software prevents tens of thousands of GPU-hours of downtime monthly, and it launched TorchSnap for checkpointing. LinkedIn, Together AI, and WhiteFiber are expanding its use.PR NewswireClockwork.io Raises $31M as LinkedIn, Together AI and WhiteFiber Adopt Its Resilience Software to Stop Wasting GPU-HoursClockwork.io raised $31 million in funding, with production deployments at LinkedIn and Together AI and expanded adoption by WhiteFiber. The company announced new TorchPass capabilities, including platform snapshots and fast application checkpoints, to preserve AI workload progress without code changes. The round brings total funding to $73 million.