AssemblyAI
AssemblyAI is a developer-focused voice AI platform offering speech-to-text, speech understanding, and bundled voice agent APIs powered by proprietary Universal models, serving AI startups, contact centers, healthcare providers, and enterprises building voice-enabled applications globally.
- Company typePrivate
- Founded2021
- HeadquartersSan Francisco, United States
- Headcount101–250
- GTM typeB2B
- OfferingSoftware
What AssemblyAI does
AssemblyAI is a San Francisco-based voice AI company founded in 2017 (originally backed by Y Combinator) that provides a developer-first API platform for converting and understanding audio at scale. The platform centers on a proprietary family of speech recognition models (Universal-2, Universal-3 Pro, Universal-3.5 Pro Realtime, and Universal-Streaming variants) that support 99+ languages, real-time streaming via WebSocket, speaker diarization with post-stream revision, mid-sentence code-switching, and an Agent Context input that reduces word error rate on real voice agent conversations to 6.99%. Around this core, AssemblyAI has assembled an integrated Voice AI platform comprising a Pre-recorded Speech-to-Text API, a Realtime Speech-to-Text API, a Speech Understanding API for entities, sentiment, topics, and summarization, a Voice Agent API that bundles STT-LLM-TTS in a single WebSocket at $4.50 per hour, an LLM Gateway that routes transcripts to GPT, Claude, Gemini, or community models with built-in fallback, and a Guardrails module for PII redaction and content moderation.
The company sells primarily through a product-led growth motion with free signup and transparent usage-based pricing ($0.45/hr for Universal-3 Pro Realtime plus optional add-ons), augmented by enterprise field sales for high-volume and regulated customers. Its target customer base spans AI startups building voice agents (Vapi, LiveKit, Pipecat, Retell AI integrations), sales and revenue intelligence platforms (Jiminny, CallRail, Delphi), contact center operators (Calabrio, EdgeTier), meeting intelligence companies (Granola, Supernormal, Grain, Dovetail), media and creator tools (Zoom, Runway, Veed.io, Happy Scribe), and healthcare providers requiring HIPAA-compliant clinical documentation. AssemblyAI reports processing 2 million hours of audio daily and supporting more than 1,000 customers as of mid-2022, with SOC 2 Type 2, HIPAA, GDPR, CCPA, and EU-U.S. Data Privacy Framework certifications in place.
AssemblyAI has raised over $113 million in venture funding across a 2017 Y Combinator pre-seed, a 2020 seed, a $28M Series A led by Accel in March 2022, a $30M Series B led by Insight Partners in July 2022, and a $50M Series C led by Accel in December 2023. Revenue mechanics are predominantly usage-based on audio duration, supplemented by annual enterprise contracts with custom pricing, volume discounts, SLAs, and self-hosted deployment for data sovereignty. The company operates from San Francisco with a corporate mailing presence in New York and supports global customers through its Voice AI Cloud and self-hosted deployment options.
AssemblyAI firmographics
Firmographics- Name
- AssemblyAI
- Legal name
- AssemblyAI, Inc.
- Website
- https://assemblyai.com
- Company type
- Private
- Founded year
- 2021
- Operating status
- Operating
- Headcount range
- 101–250 employees
- Short description
- AssemblyAI is a developer-focused voice AI platform offering speech-to-text, speech understanding, and bundled voice agent APIs powered by proprietary Universal models, serving AI startups, contact centers, healthcare providers, and enterprises building voice-enabled applications globally.
- Ownership category
- akta.pro rank
AssemblyAI industry classification
Industry- Product category
- Voice AI / Speech-to-Text API Platform
- NAICS
- Custom Computer Programming Services (541511)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- Conversational AI Platforms (chat/voice bots, orchestration) (HDAAAFAC)
- akta.pro secondary industries
- Contact Center Speech AI (agent assist, IVR automation) (HDAAAFAL), Audio & Speech Analytics (call analytics, QA, insights) (HDAAAFAH), AI Integration & Orchestration Platforms (Connectors, Workflow, iPaaS for AI) (HDAEANAI)
Keywords
Where AssemblyAI is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices2 records
Markets served
AssemblyAI business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Infrastructure, Personnel, Marketing or Sales, Operations
Revenue model
- Usage-based API Pricing: Pay-per-usage model based on audio duration processed. Universal-3 Pro Realtime at $0.45/hr, Voice Agent API bundled at $4.50/hr. No concurrency limits or throttles. Volume discounts at scale.
- Enterprise Custom Contracts: Custom pricing for high-volume enterprise customers with spend commitments and negotiated rates. Includes SLA guarantees and dedicated support.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Monthly | Free tier for initial API usage and testing |
| Usage-based | Pay-as-you-go | Universal-3 Pro Realtime at $0.45/hr base rate |
| Usage-based | Pay-as-you-go | Voice Agent API bundled pipeline at $4.50/hr flat rate |
| Usage-based | Pay-as-you-go | Pre-recorded Speech-to-Text with block-based pricing |
| Subscription | Annual | Enterprise custom pricing for high-volume deployments |
Go-to-market motion3 records
Distribution channels5 records
Marketing channels7 records
AssemblyAI product offering
Product offeringCore offering
AssemblyAI provides a Voice AI API platform that converts audio into accurate transcripts and structured insights. Its core offerings include pre-recorded and real-time Speech-to-Text APIs powered by proprietary Universal models, a Speech Understanding API for sentiment/entities/summaries, a bundled Voice Agent API that combines STT, LLM and TTS in a single WebSocket, an LLM Gateway to route transcripts to GPT/Claude/Gemini/community models, and Guardrails for PII redaction. Customers access the platform via self-serve signup, enterprise sales, orchestration platform partnerships, or self-hosted deployment.
Product overview
AssemblyAI is a Voice AI platform providing APIs for speech-to-text transcription and understanding. The core product portfolio consists of Pre-recorded Speech-to-Text API and Realtime Speech-to-Text API (powered by Universal family models including Universal-3 Pro, Universal-2, and Universal-3.5 Pro Realtime), augmented by Voice Agent API for building conversational AI agents, Speech Understanding API for extracting semantic insights, and LLM Gateway for connecting outputs to external language models. Guardrails provides PII redaction and content moderation. The platform supports both cloud-hosted (Voice AI Cloud) and self-hosted deployment options, with a no-code Playground for testing. Pricing is consumption-based with per-hour rates for streaming and free tier available.
Differentiator
Problem solved
Functional benefit
Products and services
- Pre-recorded Speech-to-Text API Converts pre-recorded audio files to text transcripts with industry-leading accuracy. Supports 99+ languages with natural language prompting for domain-specific terminology, speaker diarization, and custom vocabulary. Targeted at developers and enterprises that need batch transcription.
- Realtime Speech-to-Text API Streams real-time transcription via WebSocket with async-level accuracy and low latency. Powers voice agent applications with intelligent turn detection and sub-300ms response time. Targeted at voice AI developers building real-time conversational products.
- Voice Agent API Full-stack voice agent solution bundling speech-to-text, LLM integration, and text-to-speech in a single API at $4.50/hour flat rate. Built-in turn detection and interruption handling for natural multi-turn conversations. Targeted at developers building production voice agents.
- Speech Understanding API Extracts semantic insights from transcripts including speaker identification, sentiment analysis, chapters, summaries, topics, entities, and key phrases from a single API call. Targeted at developers and enterprises that need structured intelligence from audio.
- Guardrails Content moderation and PII redaction service that redacts sensitive personal information (SSN, credit cards, etc.) and moderates content inline on both audio and transcripts to protect privacy. Targeted at enterprises with privacy and compliance needs.
- LLM Gateway Routing layer connecting speech-to-text outputs to multiple LLM providers (GPT, Claude, Gemini, community models) from a single endpoint with built-in fallback and model swapping capabilities. Targeted at developers building generative AI features on top of audio transcripts.
- Universal-3.5 Pro Realtime Flagship real-time streaming speech-to-text model with agent context awareness, conversation memory, and voice focus features. Supports 18 languages with mid-sentence code-switching. Achieves 6.99% pooled WER on real agent conversations. Targeted at developers building high-accuracy voice agents.
- Playground No-code web interface for testing AssemblyAI's speech-to-text models with sample audio or custom uploads. Targeted at developers evaluating the API before integration.
- Self-Hosted Deployment On-premises deployment option for customers requiring data residency, private cloud hosting, or offline processing capabilities. Targeted at enterprises with data sovereignty requirements.
- Voice AI Cloud Cloud-hosted deployment with global redundancy, enterprise-grade uptime, and processing of 2 million hours of audio daily. Targeted at customers wanting a fully managed cloud solution.
Quantifiable outcome
- 6.99% word error rate on voice agent benchmark (vs 9-16% for competitors)
- +8 more outcomes
Companies that use AssemblyAI
Customer profileNamed customers20 records
Segments6 records
Ideal customer profiles3 records
AssemblyAI technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration4 records
AI capability9 records
Feature9 records
AssemblyAI partnerships and signals
Strategic signalPartnerships
Four partnerships are on record, tiered core.
- VapicoreVoice agent orchestration platform. AssemblyAI is available as a speech recognition option within Vapi's voice agent framework, enabling developers to build voice agents using AssemblyAI's STT models.
- LiveKitcoreMakes AssemblyAI Universal-3.5 Pro available through LiveKit Inference. Partnership enables real-time voice AI capabilities within LiveKit's infrastructure for voice agent applications.
- PipecatcoreVoice AI framework for building multimodal conversational agents. Uses AssemblyAI for speech recognition benchmarks. Pipecat's open STT benchmark measures Universal-3.5 Pro Realtime against competitors.
- Retell AIcorePlatform for building and deploying voice AI agents for automating real-world phone calls. Uses AssemblyAI Universal 3.5-Pro for high accuracy mode in regulated industries like healthcare and finance.
Scale indicators11 records
Recent moves6 records
Expansion highlights7 records
AssemblyAI competitors and assessment
Company assessmentDirect peers
- Rev.ai: Rev.ai provides speech-to-text APIs and asynchronous transcription services for developers and enterprises, with similar accuracy and pricing dynamics to AssemblyAI in transcription and call analytics use cases.
- Deepgram: Deepgram is a direct competitor offering speech-to-text APIs for pre-recorded and real-time transcription, with enterprise customers in contact center and conversational AI use cases — directly overlapping AssemblyAI's core STT product positioning.
- Speechmatics: Speechmatics offers an enterprise speech recognition API with broad language coverage and real-time streaming, directly competing with AssemblyAI's Universal family of models across media, contact center, and meeting intelligence workloads.
- Sonix: Sonix offers automated transcription, translation, and subtitling APIs and SaaS, serving media, podcasting, and enterprise customers with similar accuracy and language coverage ambitions to AssemblyAI's Universal-3 Pro and Universal-2 models.
- Otter.ai: Otter.ai provides AI-powered meeting transcription and conversation intelligence — overlapping with AssemblyAI's use cases in meeting notes, sales enablement, and contact center analytics, though packaged as an end-user SaaS rather than an API.
Broad incumbents
- Microsoft Azure Speech: Microsoft's Azure AI Speech offers speech-to-text and speech translation APIs within the Azure cloud ecosystem, providing an incumbent alternative for enterprises standardizing on Microsoft infrastructure for voice and transcription workloads.
- Google Cloud Speech-to-Text: Google Cloud's Speech-to-Text service is a hyperscaler-grade STT API bundled into the broader Google Cloud AI platform, competing with AssemblyAI for enterprise transcription workloads and offering comparable streaming and multilingual capabilities.
- OpenAI (Whisper): OpenAI offers Whisper and the broader speech/voice AI capabilities across its API platform. As a foundation model provider, it competes with AssemblyAI on transcription accuracy while bundling STT into a larger multimodal AI offering.
- Amazon Transcribe: AWS Transcribe provides pre-recorded and real-time speech-to-text as part of the AWS AI services portfolio, competing directly with AssemblyAI in call analytics, contact center, and media transcription use cases for AWS-native enterprises.
Emerging players
- Vapi: Vapi is a voice agent orchestration platform that lists AssemblyAI as a model option. It competes at the next layer up the stack (full voice agent framework vs. STT API) and represents both a channel partner and an emerging adjacent competitor in the voice agent infrastructure layer.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat6 records
Key risks6 records
Key highlights7 records
Customer concentration
AssemblyAI social profiles
Digital presenceAssemblyAI compliance and trust
Trust signalCompliance6 records
AssemblyAI financial estimates
Financial estimateRevenue estimate
Valuation estimate
AssemblyAI leadership team
Management profileNumber of profiles
Profiles3 records
AssemblyAI funding detail
Funding detailFunding overview
Funding rounds6 records
Investors6 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
AssemblyAI M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about AssemblyAI
What does AssemblyAI do?
AssemblyAI provides a Voice AI API platform that converts audio into accurate transcripts and structured insights. Its core offerings include pre-recorded and real-time Speech-to-Text APIs powered by proprietary Universal models, a Speech Understanding API for sentiment/entities/summaries, a bundled Voice Agent API that combines STT, LLM and TTS in a single WebSocket, an LLM Gateway to route transcripts to GPT/Claude/Gemini/community models, and Guardrails for PII redaction. Customers access the platform via self-serve signup, enterprise sales, orchestration platform partnerships, or self-hosted deployment.
Is AssemblyAI a public or private company?
AssemblyAI is a private company. It is classified as venture growth investor backed and is currently operating.
When was AssemblyAI founded?
AssemblyAI was founded in 2021. It employs 101 to 250 people.
Where is AssemblyAI based?
AssemblyAI is headquartered in San Francisco, United States, in the North America region.
How does AssemblyAI make money?
Two revenue lines are on record. Usage-based API Pricing is the primary driver. The others are enterprise Custom Contracts.
Who are AssemblyAI's main competitors?
Direct peers on record are Rev.ai, Deepgram, Speechmatics, Sonix and Otter.ai. Broad incumbents are Microsoft Azure Speech, Google Cloud Speech-to-Text, OpenAI (Whisper) and Amazon Transcribe. Vapi is listed as an emerging player.
Does AssemblyAI have an API?
Yes. AssemblyAI offers a public API providing speech-to-text and voice AI capabilities. The API supports both pre-recorded audio transcription and real-time streaming transcription via WebSocket. Authentication uses API keys. Developers can access API reference documentation, cookbooks with code examples, and streaming client libraries. The streaming API uses the assemblyai.streaming.v3 module with StreamingClient, StreamingClientOptions, StreamingEvents, and StreamingParameters classes. Developer documentation is at www.assemblyai.com/docs.
What industry is AssemblyAI in?
AssemblyAI's product category is Voice AI / Speech-to-Text API Platform. Its primary akta.pro industry code is HDAAAFAC, Conversational AI Platforms (chat/voice bots, orchestration), with a secondary code of HDAAAFAL, Contact Center Speech AI (agent assist, IVR automation). Its NAICS code is 541511 and its SIC code is 7372.