Gradium
Gradium is a Paris-based voice AI company developing unified Audio Language Models for real-time text-to-speech, speech-to-text, and voice cloning. Founded in 2025 as a Kyutai spin-out, it serves developers and enterprises across healthcare, gaming, customer support, robotics, and language services globally.
- Company typePrivate
- Founded2025
- HeadquartersParis, France
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What Gradium does
Gradium is a Paris-based voice AI company founded in September 2025 as a spin-out from the French nonprofit research lab Kyutai. Its founding team — Neil Zeghidour (CEO, formerly Meta and Google DeepMind), Olivier Teboul (CTO, formerly Google Brain), Alexandre Défossez (CSO, formerly Meta), and Laurent Mazaré (Chief Coding Officer, formerly Google DeepMind and Jane Street) — invented and open-sourced neural audio codecs and audio language models during prior research tenures, and Gradium commercializes that work as production APIs. The company emerged from stealth in December 2025 alongside a $70M seed round co-led by FirstMark Capital and Eurazeo, the largest European AI seed of that year.
Gradium's core platform is a unified Audio Language Model (ALM) family built on a proprietary Delayed Streams Modeling (DSM) architecture, delivering Text-to-Speech, Speech-to-Text, and Voice Cloning with multilingual support for English, French, Spanish, German, and Portuguese. The system is differentiated by ultra-low latency (158ms P50 Time to First Audio), independent #1 ranking on Coval TTS benchmarks against ElevenLabs, Cartesia, Deepgram, Rime, and OpenAI, voice cloning from 10-second audio samples, and a Semantic VAD that predicts turn completion from linguistic content rather than silence. Supporting products include Gradium Phonon, a ~100M-parameter on-device TTS model for Android, iOS, and MacBook CPU; Gradbot, an open-source voice-agent framework in Rust with Python bindings; and an Agent Demo at studio.gradium.ai.
Gradium operates an API-first go-to-market with a freemium-to-enterprise pricing ladder (Free, XS $13/mo, S $43/mo, M $340/mo, L $1,615/mo, plus Tailored enterprise contracts) and usage-based pay-as-you-go credits. Distribution is multi-channel: self-serve studio.gradium.ai, AWS Marketplace SaaS and Amazon SageMaker model images for regulated in-VPC deployments, direct enterprise sales, and native integrations into the LiveKit and Pipecat agent frameworks. Target segments include healthcare, gaming, contact centers, language services and media, robotics, and consumer edge devices, with named deployments at Acolad, Wonderful, InteractionLabs' Ongo robot, and the Invincible Voice ALS assistive system. A Startup Program grants six months of free M-plan to seed/Series A voice-first startups.
Gradium firmographics
Firmographics- Name
- Gradium
- Legal name
- Gradium
- Website
- https://gradium.ai
- Company type
- Private
- Founded year
- 2025
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- Gradium is a Paris-based voice AI company developing unified Audio Language Models for real-time text-to-speech, speech-to-text, and voice cloning. Founded in 2025 as a Kyutai spin-out, it serves developers and enterprises across healthcare, gaming, customer support, robotics, and language services globally.
- Ownership category
- akta.pro rank
Where Gradium is headquartered
LocationHeadquarters
- HQ city
- Paris
- HQ country
- France
- HQ region
- Europe
Offices1 record
Markets served
Gradium business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Infrastructure, Marketing or Sales, Operations
Revenue model
- API Subscription Plans: Tiered subscription model with monthly and annual billing. Plans range from Free ($0) to L ($1,615/month) with varying credit limits. Includes Studio access, API access, concurrency limits, commercial use rights, and voice cloning capabilities. Annual plans save 1 month.
- Usage-based / Pay-as-you-go: Additional credits available at per-100k rates: $6.9 (XS), $5.0 (S), $4.0 (M), $3.8 (L). Pricing becomes more favorable at higher tiers, creating volume-based economics.
- Enterprise Custom / Tailored: Custom enterprise plans with unlimited credits, unlimited instant voice clones, unlimited Pro voice clones, and custom concurrency. Contact sales for pricing.
- AWS Marketplace SaaS: Full Gradium voice AI platform available through AWS Marketplace, billed through AWS account and counting toward committed spend. L plan includes 45M credits, real-time TTS/STT, voice cloning, 15 concurrent connections.
- Amazon SageMaker Model Image: Deployable SageMaker endpoint for in-VPC inference with data residency requirements. Billed through customer's own AWS infrastructure.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free tier with basic access for evaluation |
| Subscription | Monthly | Entry-level paid plan for small projects |
| Subscription | Monthly | Popular plan for growing applications |
| Subscription | Monthly | Mid-scale production plan |
| Subscription | Monthly | Large-scale enterprise plan |
| Subscription | Multi-year contract | Custom enterprise solution |
Go-to-market motion3 records
Distribution channels5 records
Marketing channels7 records
Gradium product offering
Product offeringCore offering
Gradium develops and operates Audio Language Models sold through real-time streaming APIs for Text-to-Speech, Speech-to-Text, and Voice Cloning, plus the Gradium Studio web interface for experimentation. Customers consume credits via tiered subscriptions (Free through Tailored custom) and pay-as-you-go, with deployment available via public API, AWS Marketplace, Amazon SageMaker, or on-premises. The product supports English, French, German, Spanish, and Portuguese with sub-200ms latency and includes an on-device TTS variant (Phonon) and an open-source voice agent prototyping framework (Gradbot).
Product overview
Gradium is a voice AI company offering a unified platform of audio language models that power natural, real-time voice interactions. The core product portfolio consists of Text-to-Speech, Speech-to-Text, and Voice Cloning capabilities delivered via the Gradium API and Gradium Studio. Supporting products include Gradium Phonon (on-device TTS for edge/mobile deployment) and Gradbot (open-source prototyping framework). The platform targets voice agents across healthcare, customer support, gaming, and digital advertising, supporting English, French, German, Spanish, and Portuguese. Gradium ranks #1 on independent Coval TTS benchmarks for latency.
Differentiator
Problem solved
Functional benefit
Brands
- Phonon: On-device text-to-speech model (~100M parameters) designed for edge deployment on mobile devices, NPCs, and offline products. Runs entirely on CPU across Android and iOS with no network connection required.
- Gradbot
Products and services
- Text-to-Speech (TTS) Real-time streaming Text-to-Speech API delivering natural, expressive speech at 48kHz with word-level timestamps and multilingual support across English, French, Spanish, German, and Portuguese for developers and enterprises building voice agents.
- Speech-to-Text (STT) Real-time Speech-to-Text transcription API with controllable latency, semantic voice activity detection for natural turn-taking, robust performance in noisy environments, and code-switching support for developers and enterprises.
- Voice Cloning Voice cloning API offering instant cloning from 10 seconds of audio and Pro Voice Clones for fine-tuned models with highest market speaker similarity, integrated via cross-attention layers, for content creators, gaming, dubbing, and enterprise use cases.
- Gradium Phonon On-device Text-to-Speech model (~100M parameters) running entirely on CPU across Android and iOS with no network required, supporting voice cloning from 10-second samples and finetuned deployment for specific voice, language, and use case; designed for offline, privacy-sensitive, and high-volume consumer applications.
- Gradbot Open-source voice agent framework for prototyping voice agents in approximately 50 lines of code, built on a Rust orchestration core handling turn-taking, interruptions, silence detection, and async tool calls, integrating with LiveKit and Pipecat for production deployments.
- Gradium Studio Web-based studio for experimenting with Gradium's TTS, STT, and voice cloning models, providing real-time previews, configuration options, and API access for developers evaluating and integrating the platform.
- Gradium API WebSocket-based streaming API providing real-time Text-to-Speech, Speech-to-Text, and Voice Cloning capabilities with Python and Rust client SDKs, native LiveKit and Pipecat integrations, and enterprise options including SLAs and private cloud deployments.
- Gradium Voice on AWS Gradium voice AI platform available as a managed SaaS subscription via AWS Marketplace and as an Amazon SageMaker deployable model image for in-VPC inference, billed through AWS accounts; suited for enterprise buyers needing committed-spend billing, data residency, and HIPAA-aligned deployments.
Quantifiable outcome
- 158ms P50 Time to First Audio (TTFA) vs competitors at 294-969ms
- +3 more outcomes
Companies that use Gradium
Customer profileNamed customers5 records
Segments8 records
Ideal customer profiles3 records
Gradium technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration4 records
AI capability8 records
Feature6 records
Gradium partnerships and signals
Strategic signalPartnerships
Seven partnerships are on record, tiered core.
- AWS (Amazon)coreGradium available on AWS through two paths: (1) Managed SaaS subscription via AWS Marketplace billed through AWS account, and (2) Deployable SageMaker model image for in-VPC inference with data residency requirements.
- LiveKitcoreNative integration partnership. Gradium TTS available as LiveKit agent model for production voice agent deployments.
- PipecatcoreNative integration partnership. Gradium TTS available as Pipecat service for production voice agent deployments.
- InteractionLabs (Ongo)corePartnership to bring expressive, real-time voice AI to robotics. Ongo living lamp robot powered by Gradium's voices, Speech-to-Text and Text-to-Speech. Ongo designed with emotional presence for home robotics. Actively working together on transformational features for human-robot interaction.
- AcoladcoreStrategic partnership combining Acolad's language, data, and AI-enabled workflows with Gradium's voice AI models for real-time multilingual interpreting solutions for enterprise and public-sector use cases. Together developing next-generation AI-powered real-time interpreting solutions. Gradium also uses Acolad's data services (multilingual datasets, annotation capabilities) to train and evaluate its voice models.
- Invincible Voice (Kyutai)coreOpen-source assistive voice AI system for ALS patients and speech loss, led by Kyutai. Gradium provides core voice AI API and models including real-time speech transcription, synthesis, and voice cloning. Project is open-source and transparent.
- Wonderful (Wonderful.ai)coreGradium powers real-time voice agents on Wonderful's platform, bringing cutting-edge voice AI from experimental to deployable. Wonderful has demonstrated ability to move enterprises from pilots to reliable, operational AI.
Scale indicators6 records
Recent moves8 records
Expansion highlights7 records
Gradium competitors and assessment
Company assessmentBroad incumbents
- OpenAI: OpenAI offers TTS and STT capabilities (via Realtime API, Whisper, and TTS models) bundled into its broader AI platform. It is named in Gradium's Coval benchmark comparisons and represents a broad incumbent that can subsidize voice AI as part of a much larger model and API portfolio.
Direct peers
- Deepgram: Deepgram provides enterprise-grade speech-to-text, text-to-speech, and voice agent APIs with emphasis on accuracy and real-time performance. It is named as a direct competitor in Gradium's Coval benchmarks and serves overlapping enterprise and developer customer segments.
- Resemble AI: Resemble AI provides voice cloning, TTS, and voice authentication APIs with enterprise focus on custom voices and real-time synthesis. It is directly comparable to Gradium's voice cloning capabilities and targets similar use cases in media, gaming, and enterprise.
- Cartesia: Cartesia builds real-time audio foundation models (State Space Models) for voice agents, offering ultra-low latency TTS and STT. It competes head-to-head with Gradium on the same Coval TTS benchmark and targets the same voice agent developer market.
- Play.ht: Play.ht offers AI text-to-speech and voice cloning APIs with a large voice library and multilingual support. It competes with Gradium in the developer API segment for TTS and voice cloning, particularly for content creation and voice agents.
- ElevenLabs: ElevenLabs is the leading AI voice platform offering TTS, voice cloning, and dubbing with multilingual support. It is a direct competitor to Gradium across TTS, STT-adjacent voice agents, and voice cloning, and is named explicitly in Gradium's Coval benchmark comparisons.
- Hume AI: Hume AI builds voice and emotion AI for conversational agents with TTS, STT, and prosody/emotion recognition. It targets similar enterprise customer support and voice agent use cases as Gradium, with an emphasis on naturalness and emotional intelligence.
- WellSaid Labs: WellSaid Labs offers enterprise AI voice generation with studio-quality TTS for corporate training, e-learning, and content production. It is comparable to Gradium's enterprise TTS offering, particularly in voice quality and commercial use rights for content workflows.
- Rime: Rime provides low-latency text-to-speech APIs targeting voice agents and conversational AI applications. It is explicitly named as a competitor in Gradium's Coval benchmark comparisons and serves similar enterprise voice agent use cases.
Others
- Kyutai: Kyutai is the French nonprofit AI research lab from which Gradium spun out. It is thematically related as the source of Gradium's foundational research and talent, and continues to publish open-source audio models that indirectly compete with Gradium's proprietary stack.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key highlights7 records
Customer concentration
Gradium social profiles
Digital presenceGradium financial estimates
Financial estimateRevenue estimate
Valuation estimate
Gradium leadership team
Management profileNumber of profiles
Profiles6 records
Gradium funding detail
Funding detailFunding overview
Funding rounds2 records
Investors9 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Gradium M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Gradium
What does Gradium do?
Gradium develops and operates Audio Language Models sold through real-time streaming APIs for Text-to-Speech, Speech-to-Text, and Voice Cloning, plus the Gradium Studio web interface for experimentation. Customers consume credits via tiered subscriptions (Free through Tailored custom) and pay-as-you-go, with deployment available via public API, AWS Marketplace, Amazon SageMaker, or on-premises. The product supports English, French, German, Spanish, and Portuguese with sub-200ms latency and includes an on-device TTS variant (Phonon) and an open-source voice agent prototyping framework (Gradbot).
Is Gradium a public or private company?
Gradium is a private company. It is classified as venture growth investor backed and is currently operating.
When was Gradium founded?
Gradium was founded in 2025. It employs 1 to 10 people.
Where is Gradium based?
Gradium is headquartered in Paris, France, in the Europe region.
How does Gradium make money?
Five revenue lines are on record. API Subscription Plans are the primary driver. The others are usage-based / Pay-as-you-go, enterprise Custom / Tailored, AWS Marketplace SaaS and amazon SageMaker Model Image.
Who are Gradium's main competitors?
OpenAI is listed as a broad incumbent. Direct peers are Deepgram, Resemble AI, Cartesia, Play.ht, ElevenLabs, Hume AI, WellSaid Labs and Rime. Kyutai is listed as an others.
Does Gradium have an API?
Yes. Gradium offers WebSocket APIs designed for streaming bidirectional real-time communication. The APIs provide Text-to-Speech, Speech-to-Text, and Voice Cloning capabilities. Clients are available in Python and Rust. The API supports real-time streaming inference for voice agents with configurable parameters including delay_in_frames for transcription stability and semantic VAD for turn-taking. Enterprise plans include SLA guarantees and private cloud options for on-prem deployments. Developer documentation is at gradium.ai/api_docs.html.