Speechmatics
Speechmatics is a Cambridge, UK-headquartered voice AI company that builds proprietary speech-to-text and text-to-speech engines deployed across 55+ languages via cloud, on-premise, and on-device. It serves enterprise and developer customers in healthcare, broadcast, contact centers, and voice-agent platforms, processing 500+ years of audio monthly.
- Company typePrivate
- Founded2006
- HeadquartersCambridge, United Kingdom
- Headcount101–250
- GTM typeB2B
- OfferingSoftware
What Speechmatics does
Speechmatics is a Cambridge, UK-headquartered voice AI company that builds proprietary speech-to-text (STT) and text-to-speech (TTS) engines for enterprise and developer customers across more than 55 languages. The platform exposes a unified STT API in real-time and batch modes, with three proprietary model variants (Enhanced, Standard, and the multilingual Melia 1 with code-switching), a specialized Medical STT model (93% real-world accuracy, powered by NVIDIA infrastructure), a sub-150ms TTS API, and a Voice Agent API that combines both. The underlying models are deployable across SaaS, private cloud, container, virtual appliance, and on-device C/C++ library modes, supporting enterprise certifications (ISO/IEC 27001:2022, SOC 2 Type II, GDPR, HIPAA). The company processes 500+ years of audio per month and counts Adobe (on-device STT in Premiere Pro since 2021), AI-Media, the National Captioning Institute, NVIDIA, LiveKit, Boost.ai, VCONIC, and Sully.ai among its customers and partners.
The business operates a hybrid go-to-market: an API-first, product-led growth motion through the Speechmatics Portal (free tier with 3,000 STT minutes and 1M TTS characters per month, plus a Pro tier with usage-based pricing from $0.129/hr and a Startup Program offering up to $50,000 in credits), combined with an enterprise direct-sales motion that delivers custom contracts, dedicated Customer Success Managers and Solutions Engineers, and on-prem/on-device deployments. Revenue streams include usage-based STT and TTS billing, freemium-to-Pro conversion, enterprise licensing for on-prem and embedded deployments, bolt-on features (Translation, Summaries, Chapters, Sentiment, Topics) layered on top of base STT consumption, and OEM/embedded distribution (most notably the on-device integration inside Adobe Premiere Pro). Channel partnerships span developer platforms (LiveKit, Vapi, Pipecat, Recall.ai, Cekura, Jambonz) and strategic co-sellers in regulated industries (Boost.ai, VCONIC, AI-Media, Sully.ai).
The company was founded in 2006, is privately held, and is led by CEO Katy Wigdahl with founder Tony Robinson still associated. It has raised a total of approximately $70M across funding rounds, including a $62M Series B in June 2022 led by Susquehanna Growth Equity (with AlbionVC and IQ Capital participating). The company employs 101–250 people, holds the Queen's Award for Enterprise in Innovation (2019), and was named a G2 Leader in 2026 across multiple transcription and voice recognition categories. Recent strategic priorities (2025–2026) include the medical and bilingual Arabic-English models, the multilingual Melia model, deepening the Adobe on-device partnership, and expansion into US- and Gulf-region regulated industries via Boost.ai and Sully.ai.
Speechmatics firmographics
Firmographics- Name
- Speechmatics
- Legal name
- Speechmatics
- Website
- https://speechmatics.com
- Company type
- Private
- Founded year
- 2006
- Operating status
- Operating
- Headcount range
- 101–250 employees
- Short description
- Speechmatics is a Cambridge, UK-headquartered voice AI company that builds proprietary speech-to-text and text-to-speech engines deployed across 55+ languages via cloud, on-premise, and on-device. It serves enterprise and developer customers in healthcare, broadcast, contact centers, and voice-agent platforms, processing 500+ years of audio monthly.
- Ownership category
- akta.pro rank
Speechmatics industry classification
Industry- Product category
- Speech Recognition / Voice AI
- NAICS
- Computer Systems Design and Related Services (54151), Computer Systems Design and Related Services (5415), Translation and Interpretation Services (541930)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Programming, Data Processing, Etc. (7370), Services-Computer Integrated Systems Design (7373)
- akta.pro primary industry
- Automatic Speech Recognition (ASR) (HDAAAFAA)
- akta.pro secondary industries
- Text-to-Speech (TTS) & Voice Synthesis (HDAAAFAB), Speech Translation & Multilingual Speech Tech (HDAAAFAE), Audio & Speech Analytics (call analytics, QA, insights) (HDAAAFAH), Speech Recognition, Dictation & Medical Transcription Platforms (HLACAAAI)
Keywords
Where Speechmatics is headquartered
LocationHeadquarters
- HQ city
- Cambridge
- HQ country
- United Kingdom
- HQ region
- Europe
Offices1 record
Markets served
Speechmatics business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Marketing or Sales, Operations
Revenue model
- Usage-based Speech-to-Text API billing: Revenue from per-hour usage of batch and real-time STT models (Melia 1 $0.129/hr, Standard $0.24/hr, Enhanced $0.40/hr batch; Real-time Standard $0.24/hr, Real-time Enhanced $0.43/hr). Volume discounts auto-applied above 500 hr/month per STT type, with additional enterprise discounts starting at 24,000 hrs/year.
- Subscription tiers (Free and Pro): Free tier with monthly allowances (1,200 min real-time STT, 1,800 min batch STT, 1M TTS characters) plus Pro tier with 50 concurrent real-time sessions and higher rate limits; both Pro and Free can opt-in to Model Training in exchange for a 33% discount.
- Enterprise licensing and on-prem/on-device deployments: Enterprise customers pay custom contracts for SaaS, on-prem (container, virtual appliance), and on-device deployments including GPU/CPU models and custom voice/language development, with dedicated CSM/Solutions Engineer support.
- Text-to-Speech usage-based: Text-to-Speech billed at $0.011 per 1,000 characters on Pro tier; free allowance of 1 million characters per month; on-prem TTS deployment available for enterprise customers.
- STT bolt-on features: Per-hour usage fees for add-on capabilities: Translation $0.65/hr, Summaries $0.12/hr, Chapters $0.40/hr, Sentiment $0.12/hr, Topics $0.20/hr — attach revenue on top of base STT consumption.
- Startup Program (credits / future usage): $50,000+ in API credits offered to startup founders to drive adoption and lock in future paid usage as startups scale.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free tier — no credit card required, $0 |
| Usage-based | Pay-as-you-go | Pro tier — from $0.129/hour |
| Other | Multi-year contract | Enterprise tier — custom quote |
| Subscription | Monthly | Third-party reported starting price $0.03/month |
Go-to-market motion6 records
Distribution channels6 records
Marketing channels13 records
Speechmatics product offering
Product offeringCore offering
Speechmatics builds and sells enterprise-grade Speech-to-Text and Text-to-Speech AI engines, exposed primarily through REST APIs and consumed by developers, broadcasters, healthcare AI vendors, contact-center platforms, and OEMs. The STT engine covers 55+ languages in real-time and batch modes with built-in speaker diarization and a specialized Medical STT model, while TTS offers sub-150ms latency for voice agent use cases. Deployment options span multi-region SaaS, on-premise containers, virtual appliances, and on-device C/C++ libraries.
Product overview
Speechmatics is a unified Voice AI platform offering enterprise-grade speech recognition and synthesis, structured around a core Speech-to-Text API with complementary products. The core platform comprises the Speech-to-Text API (available in real-time and batch modes) with three proprietary models: Enhanced (highest accuracy), Standard (balanced), and Melia 1 (multilingual code-switching). Specialized models include the Medical Speech-to-Text Model (93% accuracy, healthcare-focused) and the On-Device STT model. The platform extends through the Text-to-Speech API (sub-150ms latency) and the Voice Agent API for end-to-end voice agent building. Add-on modules include Translation, and Speech Intelligence features (Summaries, Chapters, Sentiment, Topics). Speechmatics' architecture supports flexible deployment across SaaS, Private Cloud, Container, Virtual Appliance, and On-Device modes, with 55+ languages supported across transcription and 69 pairs for translation.
Differentiator
Problem solved
Functional benefit
Products and services
- Speech-to-Text API Core automatic speech recognition (ASR) API offering real-time and batch transcription across 55+ languages with three proprietary models: Enhanced (highest accuracy), Standard (balanced), and Melia 1 (multilingual code-switching). Available via SaaS, Private Cloud, Container, Virtual Appliance, and On-Device deployments for developers and enterprises integrating speech recognition into their products.
- Text-to-Speech API Low-latency Text-to-Speech API delivering sub-150ms latency, natural voices, and global scale for real-time voice agent conversations. Priced at $0.011 per 1,000 characters on Pro tier, with on-premises deployment and custom voice development available in Enterprise tier for organizations needing branded or domain-specific voices.
- Voice Agent API End-to-end Voice Agent API combining Speech-to-Text and Text-to-Speech for building responsive AI voice agents. Sub-second speaker-aware STT and TTS across 55+ languages with native integrations to agent frameworks (LiveKit, Vapi, Pipecat, Cekura, Jambonz). Targeted at developers and enterprises building conversational AI voice products.
- Medical Speech-to-Text Model Healthcare-specialized STT model achieving 93% real-world accuracy with 50% fewer errors on medical terminology and 96% medical keyword recall. Features accent-independent recognition, real-time speaker diarization, and expanded medical vocabulary including drug names, dosages, and procedures. Targeted at healthcare AI platforms, clinical documentation vendors, ambient scribes, and MedTech companies needing HIPAA-compliant medical ASR.
- Speech Intelligence AI-powered insight suite that layers Summaries, Chapters, Sentiment, and Topics on top of transcripts using LLMs, applied to live captioning, media monitoring, contact center analytics, and meeting platforms. Targeted at enterprises that need post-transcription analytics on top of base STT.
- On-Device STT On-device speech recognition model that delivers near-cloud accuracy while keeping all audio local, integrated as a C/C++ library on macOS and Windows. Processes 1 hour of audio in about 55 seconds and is licensed to OEMs (notably Adobe Premiere Pro) and enterprises needing privacy-preserving, offline-capable speech recognition.
Quantifiable outcome
- Real-time final transcripts delivered in under 1 second (60% faster than nearest competitor).
- +9 more outcomes
Companies that use Speechmatics
Customer profileNamed customers27 records
Segments9 records
Ideal customer profiles6 records
Speechmatics technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration15 records
AI capability3 records
Feature9 records
Speechmatics partnerships and signals
Strategic signalPartnerships
Ten partnerships are on record, tiered major and flagship.
- Boost.aimajorStrategic partnership to deploy enterprise-grade voice AI in Europe's most regulated industries (financial services, healthcare, public sector), combining Speechmatics STT with Boost.ai's conversational AI platform that already serves 9 of 10 Norwegian banks and 118 municipalities. Plans to extend to US market.
- VCONICmajorStrategic partnership delivering advanced conversation intelligence for healthcare and financial services. Combines VCONIC's vCon standard and Conserver platform with Speechmatics STT, achieving 93% real-world accuracy in medical terminology and real-time compliance monitoring for HIPAA, PCI-DSS, GDPR and MiFID II.
- Sully.aimajorPartnership to scale healthcare AI infrastructure globally, deploying autonomous AI agents and clinical scribes leveraging NVIDIA AI infrastructure. Combines Speechmatics' medical-grade speech models with Sully's agentic workflows — 93% real-time accuracy, 21x ROI, 30+ million minutes returned to healthcare workforce as of December 2025, with expansion into the Middle East via English-Arabic bilingual model in early 2026.
- Recall.aimajorTechnical partnership integrating Speechmatics' real-time and async transcription API natively into Recall.ai's meeting transcription platform. Users can specify 'speechmatics' as the provider to obtain high-accuracy meeting transcripts from Zoom, Google Meet, and Microsoft Teams.
- LiveKitflagshipPartnership enabling real-time, speaker-aware Voice Agents. Speechmatics brings real-time speaker diarization (built-in, not bolt-on) to LiveKit's open-source multimodal agent framework, giving 100,000+ LiveKit developers access to Speechmatics' industry-leading STT including custom dictionary (1,000 words) and 55+ languages with bilingual models.
- AI-MediaflagshipStrategic partnership to evolve captioning and language services technologies. AI-Media integrated Speechmatics' transcription engine into LEXI 3.0, enabling 120X more content delivery in 2025 than five years prior, supporting French, English, Australian Rugby League and other languages for broadcasters, governments, and education providers globally.
- AdobeflagshipAdobe became the first non-linear editing platform to include speech-to-text in Premiere in 2021, using Speechmatics on-device models. The partnership deepened in April 2026 with a new on-device STT model in Premiere Pro delivering near-cloud accuracy while keeping all audio local — processing 1 hour of audio in ~55 seconds, 12-16% better than Whisper-powered competitors, available on Windows and Mac with broad GPU support.
- NVIDIAflagshipNVIDIA powers Speechmatics' Medical Speech-to-Text model (93% accuracy). Speechmatics is part of the NVIDIA Inception Program and joined NVIDIA's Holoscan for Media partner showcase at IBC Amsterdam with Twelve Labs, Beamr, Bria, Mobius Labs, Alugha, Deepdub, Moments Lab, Qvest and Monks. Sully.ai + Speechmatics deployment also uses NVIDIA AI infrastructure.
- National Captioning Institute (NCI)majorLong-running captioning partnership — NCI chose Speechmatics' real-time ASR for its CaptionSentry service, achieving 21% usage growth in 2022 and 99% in 2023. NCI runs Speechmatics via SaaS API with engine-agnostic flexibility.
- Content GurumajorContent Guru is listed as one of Speechmatics' leading technology providers, integrating Speechmatics STT into its enterprise contact center solutions.
Scale indicators12 records
Recent moves6 records
Expansion highlights7 records
Speechmatics competitors and assessment
Company assessmentDirect peers
- Deepgram: Direct STT API competitor with real-time and batch models across 30+ languages; Speechmatics benchmarks itself directly against Deepgram (claiming 70% fewer errors at low latency and beating Deepgram on 91% of FLEURS languages via Melia). Both serve contact center, media, and voice agent developers with usage-based pricing.
- AssemblyAI: Direct STT API competitor with Universal-1 and real-time models. Speechmatics claims 50% fewer errors than AssemblyAI on real-time STT and beats AssemblyAI on 77% of FLEURS languages via Melia. Both target developers with freemium tiers and usage-based pricing, plus audio intelligence add-ons.
- Rev.ai: STT API offering streaming and asynchronous transcription with strong word-level accuracy. Overlaps with Speechmatics on meeting transcription, media, and developer API customers; smaller language coverage and no comparable on-device or medical specialization.
Broad incumbents
- OpenAI (Whisper / GPT-4o-transcribe): Open-weight and proprietary Whisper / GPT-4o-transcribe models offer broad STT that Speechmatics explicitly benchmarks against (tied on healthcare WER at 0.16). OpenAI's scale, distribution, and price-performance trajectory define the floor against which Speechmatics' accuracy premium must be defended.
- Microsoft Azure Speech: Hyperscaler STT/TTS offering that Speechmatics cites as 25% worse on real-time accuracy. Speechmatics is now also distributed via the Microsoft commercial marketplace, positioning Azure Speech both as competitor and channel.
- Google Cloud Speech-to-Text: Hyperscaler STT service bundled with broader GCP AI spend. Comparable in multi-language support and enterprise deployment options; competes head-on with Speechmatics in media captioning, contact center, and compliance workloads where Speechmatics claims accuracy and latency advantages.
- Amazon AWS Transcribe: Hyperscaler STT service bundled with AWS. Comparable on language coverage and enterprise compliance, often selected by AWS-native customers. Speechmatics' on-prem container and on-device deployments are positioned against AWS-only customers who need data residency.
- SoundHound AI: Public voice AI platform combining STT, TTS, and conversational AI for automotive, restaurant, and enterprise. Overlaps with Speechmatics in voice agent and contact center segments, with broader vertical-specific solutions and a public-market funding base.
Emerging players
- ElevenLabs: Emerging leader in voice synthesis (TTS) and increasingly voice agent infrastructure. Speechmatics' sub-150ms TTS and Voice Agent API directly overlap ElevenLabs in conversational AI stacks, while ElevenLabs does not currently match Speechmatics' enterprise STT depth.
- Gladia: Emerging STT API player tied with Speechmatics on the AIMultiple healthcare benchmark (0.16 WER) and competing in the same developer / voice agent segment. Comparable product surface (real-time, batch, diarization) but smaller language coverage and brand.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key highlights7 records
Customer concentration
Speechmatics social profiles
Digital presenceSpeechmatics compliance and trust
Trust signalCompliance5 records
Speechmatics financial estimates
Financial estimateRevenue estimate
Valuation estimate
Speechmatics leadership team
Management profileNumber of profiles
Profiles7 records
Speechmatics subsidiaries and ownership
Company hierarchySubsidiaries1 record
Speechmatics funding detail
Funding detailFunding overview
Funding rounds4 records
Investors5 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Speechmatics M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Speechmatics
What does Speechmatics do?
Speechmatics builds and sells enterprise-grade Speech-to-Text and Text-to-Speech AI engines, exposed primarily through REST APIs and consumed by developers, broadcasters, healthcare AI vendors, contact-center platforms, and OEMs. The STT engine covers 55+ languages in real-time and batch modes with built-in speaker diarization and a specialized Medical STT model, while TTS offers sub-150ms latency for voice agent use cases. Deployment options span multi-region SaaS, on-premise containers, virtual appliances, and on-device C/C++ libraries.
Is Speechmatics a public or private company?
Speechmatics is a private company. It is classified as venture growth investor backed and is currently operating.
When was Speechmatics founded?
Speechmatics was founded in 2006. It employs 101 to 250 people.
Where is Speechmatics based?
Speechmatics is headquartered in Cambridge, United Kingdom, in the Europe region.
How does Speechmatics make money?
Six revenue lines are on record. Usage-based Speech-to-Text API billing is the primary driver. The others are subscription tiers (Free and Pro), enterprise licensing and on-prem/on-device deployments, text-to-Speech usage-based, STT bolt-on features and startup Program (credits / future usage).
Who are Speechmatics's main competitors?
Direct peers on record are Deepgram, AssemblyAI and Rev.ai. Broad incumbents are OpenAI (Whisper / GPT-4o-transcribe), Microsoft Azure Speech, Google Cloud Speech-to-Text, Amazon AWS Transcribe and SoundHound AI. Emerging players are ElevenLabs and Gladia.
Does Speechmatics have an API?
Yes. Speechmatics offers a public Speech-to-Text and Text-to-Speech API for developers and enterprises. The REST-based API supports real-time and batch transcription across 55+ languages, speaker diarization, custom dictionaries, language identification, translation, summaries, sentiment, topics, and chapters. Free tier includes 3,000 minutes of STT and 1 million characters of TTS per month. Pro tier offers 50 concurrent real-time sessions and 10 file jobs per second; Enterprise tier offers unlimited scale with deployment on SaaS, Private Cloud, Container, Virtual Appliance, or On-Device (GPU and CPU). Multi-region cloud options: US, EU, Australia. Developer documentation is at docs.speechmatics.com.
What industry is Speechmatics in?
Speechmatics's product category is Speech Recognition / Voice AI. Its primary akta.pro industry code is HDAAAFAA, Automatic Speech Recognition (ASR), with a secondary code of HDAAAFAB, Text-to-Speech (TTS) & Voice Synthesis. Its NAICS code is 54151 and its SIC code is 7372.