Boson AI
Boson AI is a full-stack AI company building proprietary foundation audio models (Higgs), a real-time avatar model (Higgs Avatar), and an agentic orchestration platform (Feynman Flow) for enterprise voice AI. It serves insurance, customer support, and consumer voice experience use cases across 94 to 100+ languages.
- Company typePrivate
- Founded2023
- HeadquartersSanta Clara, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Boson AI does
Boson AI is a full-stack AI company founded in 2023 by Dr. Alex Smola and Dr. Mu Li, researchers with decades of experience building machine learning systems at scale. Headquartered in Santa Clara, California, with research and engineering operations in Toronto, Canada, and compute infrastructure hosted at the eStruxture TOR5 data center in Barrie, Ontario, the company operates as a team of roughly 30 researchers, engineers, and operators. It raised $70 million across two seed rounds in May 2023 from undisclosed major financial and strategic investors and has not publicly disclosed revenue.
The company's core products are the Higgs Audio family of foundation audio models (text-to-speech, speech-to-text, and voice chat supporting 94 to 100+ languages, built on an LLM-backbone architecture with proprietary GRPO alignment) and Feynman Flow, an agentic platform that orchestrates multi-step voice conversations and continuously improves through signals from real production calls. In May 2026 the product line expanded with Higgs Avatar v1, a real-time avatar foundation model that generates expressive face video from a single still image at roughly 16ms per frame. The technology stack is reinforced by more than 10 million hours of audio training data, more than 1 million hours of labeled STT data, a curated Voice Bank for alignment, and infrastructure partnerships with NVIDIA, eStruxture, Arc Compute, Crusoe, AWS, and Scaleway.
Boson AI generates revenue through usage-based API fees for Higgs Audio models (currently in free public preview with rate limits), enterprise subscriptions for custom voice-agent design and deployment, professional services for model customization and fine-tuning, and licensing plus enterprise support for open-source releases distributed via GitHub and Hugging Face. Distribution is multi-channel: direct enterprise field sales through Request a Demo forms, self-serve access via the Boson Workspace developer portal and Boson API, and OEM/white-label availability through Microsoft Foundry, Eigen AI, Deep Infra, and ByteCompute. Targeted customer segments span health-insurance sales, enterprise customer support, fraud detection, specialist training, AI receptionist use cases, and consumer voice experiences.
Boson AI firmographics
Firmographics- Name
- Boson AI
- Legal name
- Boson AI
- Website
- https://boson.ai
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Boson AI is a full-stack AI company building proprietary foundation audio models (Higgs), a real-time avatar model (Higgs Avatar), and an agentic orchestration platform (Feynman Flow) for enterprise voice AI. It serves insurance, customer support, and consumer voice experience use cases across 94 to 100+ languages.
- Ownership category
- akta.pro rank
Boson AI industry classification
Industry- Product category
- Voice AI / Conversational AI Foundation Models
- NAICS
- Custom Computer Programming Services (541511), Computer Systems Design Services (541512), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- Conversational AI Platforms (chat/voice bots, orchestration) (HDAAAFAC)
- akta.pro secondary industries
- Agents & Autonomous Workflows (Tool Use, Planning, Multi-Agent) (HDAAACAF), Enterprise Foundation Model Integration & APIs (Connectors, Governance, Deployment) (HDAAACAO), AI Integration & Orchestration Platforms (Connectors, Workflow, iPaaS for AI) (HDAEANAI), AI Application Enablement Platforms (Copilot/Agent Frameworks, SDKs) (HDAEANAJ), Model Serving, Inference & Deployment Platforms (APIs, Edge/On-Prem) (HDAEANAC)
Keywords
Where Boson AI is headquartered
LocationHeadquarters
- HQ city
- Santa Clara
- HQ country
- United States
- HQ region
- North America
Offices3 records
Markets served
Boson AI business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Infrastructure, Marketing or Sales, Operations
Revenue model
- API Usage Fees: Usage-based API pricing for Higgs Audio models through Boson API, currently in public preview with free access and rate limiting
- Enterprise Sales: Custom voice agent design, integration, and deployment tailored to enterprise business requirements. End-to-end pipelines customized to domain, data, and deployment
- Model Customization Services: Custom AI solutions and aligned foundation models tailored to customer needs, including fine-tuning and enterprise-specific deployments
- Open Source Models: Open source Higgs Audio 2 and Higgs-Llama models available on GitHub and Hugging Face, with enterprise support and customization as paid offerings
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Pay-as-you-go | Free public preview API access with rate limiting |
| Subscription | Annual | Enterprise custom solutions |
Go-to-market motion3 records
Distribution channels7 records
Marketing channels7 records
Boson AI product offering
Product offeringCore offering
Boson AI builds and sells proprietary voice and audio foundation AI models (the Higgs Audio family, covering TTS, STT, and voice chat across 100+ languages) along with Higgs Avatar (real-time avatar video generation) and Feynman Flow, an agentic platform that orchestrates multi-step conversational workflows for enterprise customers. Offerings are delivered via the Boson API/Boson Workspace for self-serve developers, through managed platforms such as Microsoft Foundry, Eigen AI, Deep Infra, and ByteCompute, and via custom enterprise deployments tailored to domain, data, and deployment requirements.
Product overview
Boson AI is a full-stack AI company building voice and multimodal AI systems. The core product portfolio centers on Higgs Audio (foundation audio models for TTS, STT, and voice chat), Feynman Flow (agentic orchestration platform for multi-step conversations and workflow automation), and Higgs Avatar (real-time video avatar generation). These products integrate through Boson Workspace and API, enabling developers to build real-time voice agents for enterprise workflows. Higgs Audio supports 100+ languages with voice cloning and expressive control, Feynman Flow provides continuous learning from real deployment signals, and Higgs Avatar adds visual presence to voice interactions.
Differentiator
Problem solved
Functional benefit
Products and services
- Higgs Audio Family of proprietary foundation audio models delivering production-grade text-to-speech, speech-to-text, and voice chat capabilities across 100+ TTS languages and 94 STT languages, with zero-shot voice cloning, expressive emotion and prosody control, and real-time low-latency streaming for developers and enterprises building voice agents.
- Feynman Flow Agentic platform that orchestrates real-time multi-step voice conversations, connects to business data, executes tools and workflows, and continuously improves underlying AI models through automatic fine-tuning from signals captured in real deployments, aimed at enterprise automation for customer support, sales, and training use cases.
- Higgs Avatar v1 Real-time avatar video generation model that creates live, expressive digital faces from a single still image, with lip-sync, head motion, and expression driven by audio, designed for live conversational voice agents and currently in private preview.
- Boson API REST API providing developers with access to Higgs Audio TTS, voice management, and avatar generation endpoints, including Bearer-token authentication, streaming PCM output, preset voices, reference-audio cloning, and inline emotion/style/prosody control tags, currently available in public preview with free rate-limited usage.
- Boson Workspace Self-serve developer workspace that provides API key management, a playground for Higgs Audio and Higgs Avatar, and team collaboration features as the portal through which developers and enterprises access Boson AI's voice AI offerings.
- Higgs-Llama Open-source family of large language models based on LLaMA, specially tuned for role-playing while remaining competitive in general instruction-following and reasoning, with the v2 release incorporating a Higgs Judger reward model for alignment training.
- Higgs Audio 3.0 Speech-to-Text State-of-the-art speech-to-text foundation model supporting 94 languages with language detection, sentiment and semantic understanding, and low-latency streaming transcription for enterprise and developer voice applications.
- Higgs Audio v3 TTS Chat-native text-to-speech foundation model delivering expressive conversational speech across 100+ languages, with zero-shot voice cloning and inline control over emotion, style, prosody, pauses, and sound effects for voice-agent applications.
- Higgs Audio 2.5 Voice Model Compressed voice generation model that condenses Higgs Audio to 1B parameters while improving speed (150ms time-to-first-token) and accuracy, with GRPO-aligned voice cloning and finer-grained style control over the v2 model.
Quantifiable outcome
- Higgs Audio 3.0 outperforms whisper-v3-large by large margin on key languages (English WER 1.55 vs 2.10 on librispeech_test_clean)
- +4 more outcomes
Companies that use Boson AI
Customer profileSegments6 records
Ideal customer profiles2 records
Boson AI technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration4 records
AI capability11 records
Feature8 records
Boson AI partnerships and signals
Strategic signalPartnerships
Nine partnerships are on record, tiered core, minor and supporting.
- Eigen AIcorePartnership for Bay Area Higgs Audio Hackathon (March 20-22, 2026) at Circuit Launch, Mountain View. Eigen AI hosts Higgs model suite alongside open-source LLMs and multimodal models served at scale for hackathon participants. Engineering teams will serve as judges. Exclusive access to unreleased Higgs Audio checkpoints
- Circuit LaunchminorVenue partner for Bay Area Higgs Audio Hackathon (March 20-22, 2026). Provides prototyping spaces, high-speed internet, whiteboards for collaborative development in Mountain View. Open until 12am each day of the event
- MScAC (Master of Science in Applied Computing) - University of TorontocorePartnership for Toronto Higgs Audio Hackathon (Oct 24-26, 2025) hosted at MScAC headquarters. Over 200 participants built projects using Higgs Audio TTS and ASR models. MScAC is an academic program at University of Toronto that provides industry partnerships for applied computing students
- NVIDIAsupportingTechnical support partner for infrastructure. NVIDIA GPU technology enables training and serving of Higgs audio and language models
- Arc ComputesupportingInfrastructure partner providing technical support for datacenter and cloud GPU operations
- eStruxturecoreData center partner hosting Boson AI's compute infrastructure (TOR5 facility in Barrie, Ontario) for high-performance computing and AI model training with advanced cooling systems
- CrusoesupportingCloud infrastructure partner providing GPU compute resources for model training and serving
- AWSsupportingCloud infrastructure partner providing technical support and services for cloud-based operations
- ScalewaysupportingCloud infrastructure partner providing GPU compute resources for model training and serving across European region
Scale indicators10 records
Recent moves8 records
Expansion highlights7 records
Boson AI competitors and assessment
Company assessmentDirect peers
- ElevenLabs: ElevenLabs is a leading voice AI platform offering text-to-speech, voice cloning, and dubbing APIs that directly compete with Boson AI's Higgs Audio family. Both companies target enterprise and developer use cases with multilingual, expressive speech generation.
- Hume AI: Hume AI builds emotionally intelligent voice and conversational AI, including expressive voice generation (Octave TTS) and emotion-aware speech understanding. Direct overlap with Boson's Higgs Audio v3 TTS, which emphasizes inline emotion/style control and emotional alignment.
- Deepgram: Deepgram provides enterprise speech-to-text, text-to-speech, and voice agent APIs with a developer-focused platform. Comparable to Boson's Higgs Audio 3.0 STT and Higgs Audio TTS in target market (enterprise and developers) and product category (speech AI APIs).
- AssemblyAI: AssemblyAI offers speech-to-text and speech understanding APIs (Universal-1, Universal-2) targeting developers and enterprise. Directly comparable to Boson's Higgs Audio 3.0 STT in providing multilingual, low-latency speech recognition via API.
- Speechmatics: Speechmatics delivers enterprise speech recognition with broad language coverage and accent robustness. Comparable to Boson's multilingual STT positioning and enterprise focus on high-volume transcription workloads.
- Play.ht: Play.ht provides AI text-to-speech and voice cloning APIs with a large voice library for content creators and enterprise. Overlaps directly with Boson's Higgs Audio TTS and zero-shot voice cloning capabilities.
- Resemble AI: Resemble AI offers voice cloning, text-to-speech, and real-time voice agent APIs with emphasis on enterprise security and on-premise deployment. Directly comparable to Boson's voice cloning and enterprise-focused conversational AI offering.
Broad incumbents
- OpenAI: OpenAI offers Whisper (STT), the Voice Engine / Realtime API (TTS and voice conversation), and GPT-4o multimodal as a broad AI platform with overlapping audio capabilities. OpenAI's massive scale and broader product portfolio position it as a broad incumbent in the same voice AI space.
- SoundHound AI: SoundHound provides voice AI for automotive, restaurant, and enterprise customer service with both proprietary speech recognition and conversational AI. A broader incumbent overlapping with Boson's customer support and AI receptionist use cases.
Emerging players
- Cartesia: Cartesia builds real-time multimodal AI with state-space models for low-latency voice generation and conversational agents. Comparable as an emerging voice AI startup with overlapping real-time TTS and conversational AI capabilities.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
Boson AI social profiles
Digital presenceBoson AI compliance and trust
Trust signalCompliance1 record
Boson AI financial estimates
Financial estimateRevenue estimate
Valuation estimate
Boson AI leadership team
Management profileNumber of profiles
Profiles7 records
Boson AI funding detail
Funding detailFunding overview
Funding rounds1 record
Investors1 record
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Boson AI M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Boson AI
What does Boson AI do?
Boson AI builds and sells proprietary voice and audio foundation AI models (the Higgs Audio family, covering TTS, STT, and voice chat across 100+ languages) along with Higgs Avatar (real-time avatar video generation) and Feynman Flow, an agentic platform that orchestrates multi-step conversational workflows for enterprise customers. Offerings are delivered via the Boson API/Boson Workspace for self-serve developers, through managed platforms such as Microsoft Foundry, Eigen AI, Deep Infra, and ByteCompute, and via custom enterprise deployments tailored to domain, data, and deployment requirements.
Is Boson AI a public or private company?
Boson AI is a private company. It is classified as venture growth investor backed and is currently operating.
When was Boson AI founded?
Boson AI was founded in 2023. It employs 11 to 50 people.
Where is Boson AI based?
Boson AI is headquartered in Santa Clara, United States, in the North America region.
How does Boson AI make money?
Four revenue lines are on record. API Usage Fees are the primary driver. The others are enterprise Sales, model Customization Services and open Source Models.
Who are Boson AI's main competitors?
Direct peers on record are ElevenLabs, Hume AI, Deepgram, AssemblyAI, Speechmatics, Play.ht and Resemble AI. Broad incumbents are OpenAI and SoundHound AI. Cartesia is listed as an emerging player.
Does Boson AI have an API?
Yes. Boson AI offers a REST API for voice and avatar generation. The API provides endpoints for text-to-speech (POST /v1/audio/speech), voice management (POST/GET /v1/voices), and avatar generation. Authentication uses Bearer tokens (API key format: bai-xxxx). The TTS API supports streaming PCM output, preset voices, reference audio cloning, and inline emotion/style/prosody control tags. The API is in public preview with free usage and rate limits during this phase. Available at https://api.boson.ai/ Developer documentation is at docs.boson.ai.
What industry is Boson AI in?
Boson AI's product category is Voice AI / Conversational AI Foundation Models. Its primary akta.pro industry code is HDAAAFAC, Conversational AI Platforms (chat/voice bots, orchestration), with a secondary code of HDAAACAF, Agents & Autonomous Workflows (Tool Use, Planning, Multi-Agent). Its NAICS code is 541511 and its SIC code is 7372.