Kyutai
Kyutai is a Paris-based nonprofit AI research lab founded in 2023 that builds and open-sources speech-native, multimodal, and edge AI models (Moshi, Mimi, Helium, Hibiki, Pocket TTS), funded by Iliad, CMA CGM, and Schmidt Sciences and serving AI researchers, developers, and sovereign infrastructure deployments globally.
- Company typePrivate
- Founded2023
- HeadquartersParis, France
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Kyutai does
Kyutai is a Paris-based nonprofit AI research laboratory founded in November 2023 with a mission to build and democratize artificial general intelligence through open science. The organization is funded by founding donors Iliad Group (telecommunications), CMA CGM Group (shipping and logistics), and Schmidt Sciences, and operates with approximately 20 team members comprising researchers, PhD students, postdoctoral researchers, and technical staff. All models, code, and research outputs are released publicly under open-source licenses (MIT, Apache 2.0), with no commercial pricing for its technology.
Kyutai's product portfolio centers on speech-native AI and multimodal systems. The flagship Moshi is a speech-text foundation model that processes audio directly without converting to text, achieving real-time full-duplex dialogue with theoretical 160ms latency. Supporting products include the Mimi neural audio codec (12.5Hz, 1.1kbps, jointly modeling semantic and acoustic information), Kyutai STT (streaming speech-to-text in 1B and 2.6B parameter versions), Kyutai TTS 1.6B and Pocket TTS (a 100M-parameter CPU-real-time model with voice cloning), Unmute (a modular voice AI system enabling any text LLM to listen and speak), Hibiki and Hibiki-Zero (simultaneous speech-to-speech translation with voice transfer), MoshiVis and CASA (vision-language extensions via cross-attention), and Helium-1 (a 2B-parameter multilingual LLM for edge and mobile deployment).
The organization follows a community-led go-to-market strategy through open-source distribution on GitHub and Hugging Face, supported by online demos, technical blog posts, and active social channels. There is no direct customer revenue model; downstream commercial value is captured through Gradium, a Paris-based startup spun out of Kyutai in December 2025 with $70M in seed funding, and through sovereign deployments such as the French Government's DINUM using Kyutai's subtitling technology in the Visio video conferencing platform deployed to 200,000 public agents.
Kyutai firmographics
Firmographics- Name
- Kyutai
- Legal name
- Kyutai
- Website
- https://kyutai.org
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Kyutai is a Paris-based nonprofit AI research lab founded in 2023 that builds and open-sources speech-native, multimodal, and edge AI models (Moshi, Mimi, Helium, Hibiki, Pocket TTS), funded by Iliad, CMA CGM, and Schmidt Sciences and serving AI researchers, developers, and sovereign infrastructure deployments globally.
- Ownership category
- akta.pro rank
Kyutai industry classification
Industry- Product category
- Speech AI Software
- NAICS
- Computer Systems Design and Related Services (5415)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- Conversational AI Platforms (chat/voice bots, orchestration) (HDAAAFAC)
- akta.pro secondary industries
- Speech Translation & Multilingual Speech Tech (HDAAAFAE), Enterprise Foundation Model Integration & APIs (Connectors, Governance, Deployment) (HDAAACAO)
Keywords
Where Kyutai is headquartered
LocationHeadquarters
- HQ city
- Paris
- HQ country
- France
- HQ region
- Europe
Offices1 record
Markets served
Kyutai business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Infrastructure, Operations
Revenue model
- Research Funding and Grants: Kyutai is a non-profit research lab funded by Iliad Group, CMA CGM Group, and Schmidt Sciences. They operate as an open-science organization without a commercial revenue model. All models and research are released publicly under open-source licenses.
Go-to-market motion1 record
Distribution channels3 records
Marketing channels7 records
Kyutai product offering
Product offeringCore offering
Kyutai is a Paris-based non-profit AI research lab that develops and publicly releases open-source speech and multimodal AI models. Its core offerings include speech-native dialogue systems (Moshi), streaming text-to-speech (Kyutai TTS 1.6B, Pocket TTS), streaming speech-to-text (Kyutai STT), modular voice AI (Unmute), simultaneous speech-to-speech translation (Hibiki, Hibiki-Zero), and vision-language models (MoshiVis, CASA). All models, code, and research are released under open-source licenses for free commercial and research use.
Product overview
Kyutai is a Paris-based nonprofit AI research lab dedicated to open science, building and democratizing artificial general intelligence. The product portfolio centers on speech-native AI technology, with Moshi as the flagship speech-text foundation model enabling real-time full-duplex dialogue. Supporting products include Kyutai STT (streaming speech-to-text) and Kyutai TTS 1.6B/Pocket TTS (text-to-speech models), unified through the Unmute modular voice AI system. Vision capabilities are added via MoshiVis using cross-attention adaptation, while Hibiki/Hibiki-Zero provide simultaneous speech-to-speech translation. The Mimi neural audio codec serves as foundational infrastructure across all audio products. Language model capabilities are provided by Helium-1 (2B parameter multilingual LLM). All products are released as open-source code and model weights for reproducibility.
Differentiator
Problem solved
Functional benefit
Products and services
- Moshi The first speech-native dialogue system that processes speech directly rather than converting to text and back, enabling minimal latency and understanding of emotions and non-verbal aspects of communication. Uses multi-stream architecture modeling user and model audio in parallel.
- Hibiki Simultaneous speech-to-speech translation model for French-to-English translation. Transfers speaker's voice and flow with quality closest to human interpreters. 2B parameter model with 1B mobile version for on-device inference.
- Hibiki-Zero Simultaneous speech-to-speech translation model supporting French, Spanish, Portuguese, and German to English. Uses reinforcement learning eliminating need for word-level alignment data. 3B parameter decoder-only model.
- Unmute Modular voice AI system that enables any LLM to listen and speak using Kyutai's speech-to-text and text-to-speech models. Response latency below one second. Supports function calling and external tool integration.
- Helium-1 2B parameter modular and multilingual language model supporting 6 languages (English, French, German, Italian, Portuguese, Spanish). Designed for edge and mobile deployment with focus on latency and privacy.
- Pocket TTS 100M parameter text-to-speech model with voice cloning capability. Small enough to run in real-time on CPU. Achieves lowest word error rate (1.84) among comparable models while being the only one faster than real-time on CPU.
- Kyutai TTS 1.6B Streaming text-to-speech model with 1.6B parameters based on delayed streams modeling. Used in Unmute for low-latency applications such as voice assistants. Supports multiple voices across US, UK, France, and Ireland.
- Kyutai STT Streaming speech-to-text model optimized for real-time usage with semantic voice activity detection. Available in 1B (English/French) and 2.6B (English-only) versions. Supports PyTorch, Rust, and MLX implementations.
- MoshiVis Vision Speech Model extending Moshi to discuss images while preserving real-time latency and natural conversation abilities. Uses lightweight cross-attention adaptation modules trained on frozen Moshi backbone.
- Mimi Neural Audio Codec Streaming neural audio codec that efficiently models both semantic and acoustic information at 12.5Hz and 1.1kbps while achieving real-time latency. Jointly models semantic and acoustic information using distillation from WavLM.
- CASA Cross-Attention over Self-Attention framework for efficient vision-language fusion. Allows adapting pretrained token-insertion VLMs to enjoy practical benefits of cross-attention with near-constant memory cost and inference cost over video streams.
Quantifiable outcome
- First real-time full-duplex spoken large language model with 160ms theoretical latency (200ms in practice)
- +2 more outcomes
Companies that use Kyutai
Customer profileNamed customers1 record
Segments3 records
Ideal customer profiles3 records
Kyutai technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability10 records
Feature9 records
Kyutai partnerships and signals
Strategic signalPartnerships
Two partnerships are on record, tiered flagship and core.
- GradiumflagshipGradium was spun out from Kyutai to commercialize ultra-low latency multilingual voice AI technology. Kyutai's research forms the foundation for Gradium's products. Gradium raised $70M in seed funding from FirstMark Capital, Eurazeo, DST Global Partners, and others.
- French Government (DINUM)coreThe French government's Interministerial Digital Directorate (DINUM) uses Kyutai's subtitling technology in their sovereign Visio platform, deployed to 200,000 public agents. This represents production deployment of Kyutai's research in national AI infrastructure.
Scale indicators4 records
Recent moves7 records
Expansion highlights6 records
Kyutai competitors and assessment
Company assessmentDirect peers
- Sesame: Sesame builds a real-time conversational voice assistant (CSM) and explicitly uses Kyutai's Mimi neural audio codec. It directly competes in full-duplex spoken dialogue AI, the same core product category as Moshi.
- Cartesia: Cartesia develops real-time, low-latency voice AI models (state-space models) for developers and enterprises, competing directly with Kyutai's Unmute and Kyutai TTS offerings in the streaming voice-AI stack.
- Hume AI: Hume AI builds an emotionally aware voice and conversational AI platform (EVI) targeting natural, full-duplex spoken interaction — overlapping closely with Kyutai's Moshi and emotion-aware speech-native positioning.
- PlayHT: PlayHT provides real-time TTS and voice cloning APIs targeting developers and enterprises, directly comparable to Kyutai's Pocket TTS and Kyutai TTS products on the same core capability set.
Broad incumbents
- OpenAI: OpenAI ships GPT-4o with real-time voice mode, the most prominent commercial full-duplex spoken LLM. It is a broad incumbent defining the category Kyutai's Moshi targets, with vastly larger distribution and capital.
- Google DeepMind: Google DeepMind develops Gemini with voice capabilities and publishes foundational speech research, competing with Kyutai on both end products and underlying open speech-translation/speech-LLM research.
- Meta FAIR: Meta FAIR released SeamlessM4T and other open speech-to-speech/translation models, the closest open-source incumbent to Kyutai's Hibiki/Hibiki-Zero simultaneous translation research and a frequent benchmark reference.
- ElevenLabs: ElevenLabs is a leading commercial voice-AI platform for TTS, voice cloning, and conversational agents — broadly overlapping with Kyutai's Pocket TTS, Kyutai TTS 1.6B, and Unmute offerings at production scale.
Emerging players
- Mistral AI: Mistral AI is a Paris-based open-weight foundation model lab — the closest European geographic and open-science peer to Kyutai, and a potential partner/competitor for sovereign AI and open-LLM efforts.
- Rime: Rime builds production-grade TTS and voice-cloning APIs for enterprise, addressing the same developer and enterprise voice-AI customer base that Unmute and Kyutai TTS serve.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat5 records
Key risks5 records
Key highlights6 records
Customer concentration
Kyutai social profiles
Digital presenceKyutai financial estimates
Financial estimateRevenue estimate
Valuation estimate
Kyutai leadership team
Management profileNumber of profiles
Profiles9 records
Kyutai funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Kyutai M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Kyutai
What does Kyutai do?
Kyutai is a Paris-based non-profit AI research lab that develops and publicly releases open-source speech and multimodal AI models. Its core offerings include speech-native dialogue systems (Moshi), streaming text-to-speech (Kyutai TTS 1.6B, Pocket TTS), streaming speech-to-text (Kyutai STT), modular voice AI (Unmute), simultaneous speech-to-speech translation (Hibiki, Hibiki-Zero), and vision-language models (MoshiVis, CASA). All models, code, and research are released under open-source licenses for free commercial and research use.
Is Kyutai a public or private company?
Kyutai is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was Kyutai founded?
Kyutai was founded in 2023. It employs 11 to 50 people.
Where is Kyutai based?
Kyutai is headquartered in Paris, France, in the Europe region.
How does Kyutai make money?
One revenue line is on record: research Funding and Grants.
Who are Kyutai's main competitors?
Direct peers on record are Sesame, Cartesia, Hume AI and PlayHT. Broad incumbents are OpenAI, Google DeepMind, Meta FAIR and ElevenLabs. Emerging players are Mistral AI and Rime.
Does Kyutai have an API?
No public API is recorded for Kyutai.
What industry is Kyutai in?
Kyutai's product category is Speech AI Software. Its primary akta.pro industry code is HDAAAFAC, Conversational AI Platforms (chat/voice bots, orchestration), with a secondary code of HDAAAFAE, Speech Translation & Multilingual Speech Tech. Its NAICS code is 5415 and its SIC code is 7372.