Mozilla’s Common Voice
Mozilla Common Voice is an open-source, crowdsourced multilingual speech dataset initiative operated by the Mozilla Foundation, serving ASR researchers, developers, and technology firms by providing free Creative Commons-licensed voice data across 100+ languages, including under-resourced locales.
- Company typePrivate
- Founded2017
- HeadquartersSan Francisco, United States
- Headcount—
- GTM typeB2B
- OfferingSoftware
What Mozilla’s Common Voice does
Mozilla Common Voice is an open-source, crowdsourced multilingual speech recognition dataset platform operated by the Mozilla Foundation, a 501(c)(3) nonprofit based in San Francisco. Founded in 2017, the initiative enables volunteers worldwide to contribute voice recordings, validate other contributors' clips, and supply demographic metadata, building a publicly available corpus that spans 100+ languages — including many under-resourced languages systematically absent from commercial speech datasets. The platform's technical architecture consists of a web-based contribution interface, a validation pipeline, and a release process that publishes audio and transcripts under a Creative Commons license, with structured metadata (age, gender, accent) attached to each clip.
The platform serves speech and language AI researchers, academic labs, startups, and large technology firms that need training data for automatic speech recognition (ASR) and conversational AI systems, particularly in languages underserved by commercial cloud providers. Notable product additions include Spontaneous Speech Beta (2024), which extends coverage beyond scripted read-aloud recordings into conversational speech, and the African Next Voices project (2025), a Gates Foundation- and Meta-funded expansion focused on Sub-Saharan African languages.
Mozilla Common Voice does not operate a commercial revenue model: the dataset is distributed free of charge under Creative Commons terms, and the initiative is sustained through philanthropic grants and Mozilla Foundation sponsorship rather than product pricing, subscriptions, or licensing fees. The go-to-market is community- and ecosystem-driven — researchers, developers, and partner organizations discover and adopt the dataset through open-source channels, academic citations, and integrations into ASR frameworks — with no sales force, enterprise tier, or paid product lines.
Mozilla’s Common Voice firmographics
Firmographics- Name
- Mozilla’s Common Voice
- Legal name
- Mozilla Foundation
- Website
- https://commonvoice.mozilla.org
- Company type
- Private
- Founded year
- 2017
- Operating status
- Operating
- Short description
- Mozilla Common Voice is an open-source, crowdsourced multilingual speech dataset initiative operated by the Mozilla Foundation, serving ASR researchers, developers, and technology firms by providing free Creative Commons-licensed voice data across 100+ languages, including under-resourced locales.
- Ownership category
- akta.pro rank
Mozilla’s Common Voice industry classification
Industry- Product category
- Open-Source Speech Recognition Dataset
- NAICS
- Other Sound Recording Industries (512290)
- akta.pro primary industry
- Speech Translation & Multilingual Speech Tech (HDAAAFAE)
Keywords
Where Mozilla’s Common Voice is headquartered
LocationHeadquarters
- HQ city
- San Francisco
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Mozilla’s Common Voice business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Infrastructure, Operations
Revenue model
- Non-Commercial Open Dataset Distribution: Mozilla Common Voice is a non-profit initiative of the Mozilla Foundation. The dataset is freely distributed under Creative Commons licensing to anyone who wants to use it for building speech recognition models. There is no commercial revenue generation; the initiative is funded through Mozilla Foundation's broader nonprofit operations and grants.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Pay-as-you-go | Free open access to the complete Common Voice dataset and platform tools |
Go-to-market motion1 record
Distribution channels3 records
Marketing channels5 records
Mozilla’s Common Voice product offering
Product offeringCore offering
Mozilla's Common Voice is an open-source, crowdsourced speech dataset platform that collects voice recordings and transcriptions across 100+ languages. Community members contribute via reading sentences, answering spontaneous speech prompts, validating recordings, and writing new sentences, with the resulting dataset distributed freely under Creative Commons licensing for training AI speech recognition models.
Product overview
Mozilla's Common Voice is a single unified crowdsourced speech data platform with modular contribution and validation components. The platform operates as a dataset collection system where users contribute via Speak (reading sentences) or Spontaneous Speech (answering questions), validate via Listen (validating readings) or Transcribe Audio, and write via Write (adding sentences). The collected and validated voice recordings are made available as a downloadable dataset for training AI speech recognition models. The platform supports multiple languages and includes a dedicated section for community language management.
Differentiator
Problem solved
Functional benefit
Brands
- Spontaneous Speech Beta: Beta feature for collecting unscripted, natural speech patterns from contributors for voice recognition training data.
Products and services
- Common Voice Platform A crowdsourced voice dataset platform that enables users to contribute voice recordings (via Speak and Spontaneous Speech), validate audio (via Listen), contribute text (via Write), and download speech recordings for training AI speech recognition models across 100+ languages. Designed for AI/ML researchers, developers, academic institutions, and companies building voice-enabled applications.
- Common Voice Dataset Download A downloadable collection of curated voice recordings and associated metadata made available to researchers, developers, and companies for training speech recognition models. Distributed under Creative Commons licensing at no cost, supporting 100+ languages.
Quantifiable outcome
- Over 100 languages supported with crowdsourced contributions from global community
- +1 more outcomes
Companies that use Mozilla’s Common Voice
Customer profileNamed customers3 records
Segments3 records
Ideal customer profiles2 records
Mozilla’s Common Voice technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability2 records
Feature3 records
Mozilla’s Common Voice partnerships and signals
Strategic signalPartnerships
One partnership is on record.
- African Next Voices (Gates Foundation & Meta)majorMozilla Common Voice is a key data partner in the African Next Voices initiative, a collaborative effort funded by the Gates Foundation and Meta to create the largest dataset of African languages for AI development. The project aims to improve AI understanding and processing of African languages, which are currently underrepresented in AI tools. Common Voice contributes its crowdsourced data collection infrastructure and language coverage expertise to this initiative.
Scale indicators3 records
Recent moves5 records
Expansion highlights4 records
Mozilla’s Common Voice competitors and assessment
Company assessmentDirect peers
- LibriSpeech: LibriSpeech is one of the most widely used open speech corpora for English ASR research, derived from public-domain audiobooks. It is a direct functional peer to Common Voice for ASR training, though narrower in language coverage (English-only).
- VoxPopuli: VoxPopuli is a large-scale multilingual speech corpus for European languages released under a permissive license. It competes with Common Voice in the multilingual open speech dataset category, with backing from Meta/Facebook Research.
- OpenSLR (Open Speech and Language Resources): OpenSLR hosts a large collection of open speech and language datasets for academic and commercial use. It is the closest comparison to Mozilla Common Voice as a free, open-source speech data repository serving the same developer and researcher audience with multilingual corpora.
Others
- Hugging Face: Hugging Face hosts Common Voice datasets and many competing open speech corpora on its platform. It is an enabling ecosystem participant and indirect distribution channel for Common Voice, as well as a venue where competing open datasets gain developer traction.
- EleutherAI: EleutherAI is a nonprofit AI research lab that curates and releases open datasets for language model training. It is an adjacent peer in the open-data-for-AI ecosystem, sharing Common Voice's nonprofit governance model and open-data philosophy, though focused on text rather than speech.
Regional players
- AISHELL Foundation: AISHELL publishes large open Mandarin speech corpora widely used in Chinese ASR research. It is comparable to Common Voice as a major open speech dataset, but is regionally focused on Chinese rather than globally multilingual.
Broad incumbents
- Rev AI (Rev.com): Rev is a commercial transcription and ASR provider that maintains a large proprietary speech corpus from its human transcription business. While a commercial alternative, it serves the same downstream ASR market and competes for the speech data value chain that Common Voice occupies for open use.
- Mozilla DeepSpeech: Mozilla DeepSpeech is an open-source ASR engine that historically relied on Common Voice training data. It is a sibling Mozilla project and a direct consumer of Common Voice, illustrating the technology stack relationship and the Mozilla ecosystem around open speech.
- Meta AI (Omnilingual ASR): Meta's Omnilingual ASR initiative covers 1,600+ languages with proprietary datasets and is a broad strategic competitor in the speech data space. While Meta is also a partner through African Next Voices, its proprietary data efforts compete with Common Voice for both researchers and downstream model quality.
Emerging players
- Common Voice (FLEURS / Mozilla sister projects): FLEURS is a multilingual speech dataset from Google Research covering 100+ languages. It overlaps directly with Common Voice in language coverage and is a peer for academic and developer users benchmarking low-resource ASR.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat4 records
Key risks6 records
Key highlights6 records
Customer concentration
Mozilla’s Common Voice social profiles
Digital presenceMozilla’s Common Voice financial estimates
Financial estimateRevenue estimate
Valuation estimate
Mozilla’s Common Voice leadership team
Management profileNumber of profiles
Mozilla’s Common Voice funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Mozilla’s Common Voice M&A and investment
M&A and investmentM&A
Investments7 records
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Mozilla’s Common Voice
What does Mozilla’s Common Voice do?
Mozilla's Common Voice is an open-source, crowdsourced speech dataset platform that collects voice recordings and transcriptions across 100+ languages. Community members contribute via reading sentences, answering spontaneous speech prompts, validating recordings, and writing new sentences, with the resulting dataset distributed freely under Creative Commons licensing for training AI speech recognition models.
Is Mozilla’s Common Voice a public or private company?
Mozilla’s Common Voice is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was Mozilla’s Common Voice founded?
Mozilla’s Common Voice was founded in 2017.
Where is Mozilla’s Common Voice based?
Mozilla’s Common Voice is headquartered in San Francisco, United States, in the North America region.
How does Mozilla’s Common Voice make money?
One revenue line is on record: non-Commercial Open Dataset Distribution.
Who are Mozilla’s Common Voice's main competitors?
Direct peers on record are LibriSpeech, VoxPopuli and OpenSLR (Open Speech and Language Resources). Others are Hugging Face and EleutherAI. AISHELL Foundation is listed as a regional player. Broad incumbents are Rev AI (Rev.com), Mozilla DeepSpeech and Meta AI (Omnilingual ASR). Common Voice (FLEURS / Mozilla sister projects) is listed as an emerging player.
Does Mozilla’s Common Voice have an API?
No public API is recorded for Mozilla’s Common Voice.
What industry is Mozilla’s Common Voice in?
Mozilla’s Common Voice's product category is Open-Source Speech Recognition Dataset. Its primary akta.pro industry code is HDAAAFAE, Speech Translation & Multilingual Speech Tech. Its NAICS code is 512290.