Mozilla Data Collective
Mozilla Data Collective operates a data sharing platform and marketplace for AI training data, enabling researchers, NGOs, and language communities to publish and license datasets while retaining ownership. A 5% transaction fee on compensated datasets funds the mission-locked, Mozilla Foundation-incubated entity.
- Company typePrivate
- Founded2025
- HeadquartersLondon, United Kingdom
- Headcount—
- GTM typeB2B
- OfferingDigital Commerce or Content
What Mozilla Data Collective does
Mozilla Data Collective (MDC) operates a data sharing platform and marketplace for artificial intelligence training data, founded in September 2025 as a mission-locked British company (company number 17054959) and incubated by Mozilla Foundation. The platform enables data contributors—including researchers, NGOs, language communities, broadcasters, and civic organizations—to publish, license, and distribute datasets while retaining ownership and control over terms, with three access models available: openly licensed (CC0, CC-BY), community-governed, or compensated. MDC serves two participant groups: data contributors seeking fair value exchange and provenance preservation, and data consumers (AI/ML researchers, model trainers, technical organizations) requiring ethically-sourced, diverse training data for underserved languages and modalities.
The platform is built on cloud storage infrastructure (S3/R2) with a REST API, Python SDK, and presigned URLs for programmatic dataset access, supporting resumable downloads via range requests and enforcing 30 downloads/day per organization rate limits. Supporting features include an AI-powered Data Assistant for natural-language dataset discovery (Alpha, launched May 2026), a Request to Access gating workflow with Public/Restricted/Private options, a Data Provider Analytics Portal, a Dataset Notification subscription feature, and an open-source Apache 2.0 speech-to-text fine-tuning blueprint for Whisper and MMS models. MDC also operates a competition platform powered by DrivenData, hosting the inaugural Lost In Transcription competition with $20,000 in prizes for speech recognition models on underrepresented languages (Indonesian-Javanese, Nahuatl-Spanish, Spanish-English). Dataset coverage spans 1,001+ assets across ASR, TTS, MT, NLP, CV, and multimodal modalities, with global geographic reach including African, South Asian, European, and Latin American language communities.
MDC's business model is a marketplace take rate: data uploaders set their own prices for compensated datasets (14 listings among 1,001+ total), and MDC retains 5% per transaction to cover infrastructure costs while 100% of license fees flow to contributors. Payment processing is handled by Stripe. The go-to-market is community-led, leveraging Mozilla Foundation's existing networks, developer relations, conference presence (PyConAU 2025, Mozilla Festival Barcelona), and AI model competitions to attract both contributors and consumers. Distribution is primarily self-serve via the web platform, REST API, and a dedicated competitions site. No external funding rounds are disclosed; the company operates under Mozilla Foundation incubation with stated ambition for self-sustainability through platform fees rather than grant dependence.
Mozilla Data Collective firmographics
Firmographics- Name
- Mozilla Data Collective
- Legal name
- Mozilla Data Collective LTD
- Website
- https://mozilladatacollective.com
- Company type
- Private
- Founded year
- 2025
- Operating status
- Operating
- Short description
- Mozilla Data Collective operates a data sharing platform and marketplace for AI training data, enabling researchers, NGOs, and language communities to publish and license datasets while retaining ownership. A 5% transaction fee on compensated datasets funds the mission-locked, Mozilla Foundation-incubated entity.
- Ownership category
- akta.pro rank
Mozilla Data Collective industry classification
Industry- Product category
- AI Data Marketplace
- NAICS
- Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (51821)
- SIC
- Services-Computer Programming, Data Processing, Etc. (7370)
- akta.pro primary industry
- Dataset Discovery, Marketplaces & Licensing (HDAAALAE)
- akta.pro secondary industries
- Data Sharing, Data Exchange & Data Marketplace Platforms (HDAEABAH), Data Pipelines for GenAI (Curation, Filtering, Deduplication, Copyright) (HDAAACAJ), Data Sharing & Data Marketplace Governance (HDAEADAK)
Keywords
Where Mozilla Data Collective is headquartered
LocationHeadquarters
- HQ city
- London
- HQ country
- United Kingdom
- HQ region
- Europe
Offices1 record
Markets served
Mozilla Data Collective business model
Business model- GTM type
- B2B
- Offering type
- Digital Commerce or Content
- Cost components
- Technology or R&D, Infrastructure, Operations, Personnel, Marketing or Sales
Revenue model
- Platform Fee on Dataset Transactions: MDC charges a 5% platform service fee on compensated dataset purchases, bundled into the total price shown to data consumers at checkout. This covers infrastructure, hosting, and support costs. The remaining 95% of dataset license fees goes directly to the data uploader.
- Dataset Licensing Fees: Data providers can set prices for their datasets. All dataset fees (100%) go directly to uploaders. MDC acts as a payment facilitation intermediary, processing payments through Stripe. This revenue stream supports uploaders including communities with lived experience of colonialism who need value recognition for their labor.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Freemium | Pay-as-you-go | Open/Community Datasets - Free access |
| Transaction based/ take rate | Pay-as-you-go | Compensated Datasets - Variable pricing set by uploaders |
Go-to-market motion3 records
Distribution channels4 records
Marketing channels7 records
Mozilla Data Collective product offering
Product offeringCore offering
Mozilla Data Collective operates a cloud-based data sharing platform and marketplace for AI training datasets, enabling data contributors to upload, license, and distribute datasets under open, community-governed, or compensated models. Data consumers—including AI/ML researchers, developers, and organizations—discover and download multilingual, multicultural, and multimodal datasets through a web interface, REST API, and Python SDK. The platform charges a 5% fee on compensated dataset transactions, with 100% of dataset license fees going to contributors.
Product overview
Mozilla Data Collective is a data sharing platform and marketplace built around human data agency and fair value exchange. The core platform (mdc-platform) enables data uploaders to publish, license, and distribute datasets with full ownership and control, while data consumers discover and download datasets for AI/ML training. The platform supports three access models: openly licensed (CC0, CC-BY), community-governed, and compensated (paid access with MDC taking a 5% platform fee). Supporting products include a REST API and Python SDK for programmatic access, the speech-to-text-finetune open-source blueprint for Whisper/MMS model fine-tuning, a Data Assistant for AI-powered dataset search, a Data Provider Analytics Portal, an MDC Competitions platform powered by DrivenData, a Request to Access gating feature, and a Dataset Notification subscription feature.
Differentiator
Problem solved
Functional benefit
Products and services
- Mozilla Data Collective Platform Core cloud-based data sharing platform and marketplace enabling data uploaders to publish, license, and distribute AI training datasets while retaining ownership and control. Data consumers discover and download multilingual, multicultural, and multimodal datasets through a web interface. Supports three access models: openly licensed (CC0, CC-BY), community-governed, and compensated with a 5% platform fee.
- MDC Public API REST API enabling developers to programmatically discover, access, and download datasets from the MDC platform in any programming language. Base URL: dev.mozilladatacollective.com/api. Supports GET /datasets/:datasetId and POST /datasets/:datasetId/download endpoints with presigned URL downloads.
- MDC Python SDK Python software development kit providing convenient access to MDC platform datasets. Enables dataset loading, authentication, and download management for Python-based machine learning workflows.
- speech-to-text-finetune Blueprint Open-source blueprint enabling developers to fine-tune Whisper-based or MMS-based speech-to-text models using MDC-hosted datasets. Includes local transcription app, model comparison demo, and GitHub Codespaces/Google Colab instant run options.
- MDC Competitions AI model competition platform powered by DrivenData, hosting prize-based competitions for building models on underserved linguistic datasets. Currently features the Lost In Transcription competition for speech recognition with $20,000 in prizes across Indonesian-Javanese, Nahuatl-Spanish, and Spanish-English contexts.
Quantifiable outcome
- 100% of dataset fees go directly to data contributors
- +2 more outcomes
Companies that use Mozilla Data Collective
Customer profileNamed customers6 records
Segments4 records
Ideal customer profiles4 records
Mozilla Data Collective technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration2 records
AI capability9 records
Feature6 records
Mozilla Data Collective partnerships and signals
Strategic signalPartnerships
Ten partnerships are on record, tiered core and minor.
- Mozilla FoundationcoreMozilla Data Collective is incubated by Mozilla Foundation and backed by Mozilla.org. MDC is a mission-locked British company governed by Mozilla Foundation to ensure the platform's purpose is preserved. This relationship provides MDC with Mozilla's community networks, brand credibility, and organizational infrastructure.
- CLARIN (Common Language Resources and Technology Infrastructure)minorMozilla Data Collective datasets are now discoverable through CLARIN's Virtual Language Observatory, expanding visibility for community-governed language datasets and improving exploration for European researchers, developers, and language technology practitioners.
- DrivenDatacoreMDC Competitions platform is powered by DrivenData (competitions.mozilladatacollective.com). This provides the technical infrastructure for running AI model competitions with prize components.
- Common VoicecoreMozilla Data Collective hosts segments and derivatives of Mozilla Common Voice datasets (CC0-licensed speech data). Common Voice launched at Mozilla Foundation in 2017. MDC provides stewardship support including respecting right to be forgotten and maintaining versioning.
- Institute of African Digital HumanitiesminorContributes African language datasets including TTS datasets, sociocultural datasets, and ALCAM multimodal datasets for underrepresented African languages (Bafia, Mada, Suundi, Foulbe, etc.).
- AmnesiaminorMexican civil society organization preserving public information. Contributes 40+ datasets documenting freedom of information filings across Mexican states, addressing digital amnesia caused by government transitions.
- PROXIMAminorContributes Urdu and Sindhi NLP datasets for LLM training including instruction tuning datasets, synthetic datasets, and specialized domain datasets (medical, legal, financial).
- Radio Free Europe/Radio Liberty (RFE/RL)minorInternational broadcaster reaching 44 million people weekly across 18 countries in 24 languages, contributing datasets to serve communities through AI-ready language data.
- Open Home FoundationminorContributes TTS datasets to enable voice assistants to speak underserved languages, supporting open source voice assistant development beyond mainstream languages.
- StripecorePayment processor for dataset transactions. Stripe facilitates distribution of dataset fees to uploaders and platform fees to MDC, enabling the compensated datasets feature.
Scale indicators4 records
Recent moves8 records
Expansion highlights6 records
Mozilla Data Collective competitors and assessment
Company assessmentDirect peers
- Hugging Face: Hugging Face operates the largest community hub for AI datasets and models, including a Datasets library and Spaces for dataset hosting — directly overlapping MDC's dataset discovery and hosting core.
- Kaggle: Kaggle Datasets is a free dataset marketplace widely used by AI/ML practitioners; serves the same dataset discovery and consumption workflow as MDC, with a much larger user base.
- Roboflow: Roboflow hosts a large catalog of computer vision datasets with discovery and tooling; overlaps MDC's CV dataset hosting and developer-focused dataset distribution.
Broad incumbents
- AWS Data Exchange: AWS Data Exchange is a cloud-native marketplace for third-party data, including AI training datasets; competes for enterprise buyers seeking licensed datasets with managed billing and distribution.
- Scale AI: Scale AI is a major AI data platform spanning labeling and dataset curation across modalities; competes for the same AI training data customers MDC serves, with significant enterprise sales reach.
- Appen: Appen is an established AI training data provider with global crowd contributors and multilingual speech/language datasets; competes for the same model-training data sourcing budgets.
Emerging players
- Common Crawl: Common Crawl is a nonprofit providing openly licensed web crawl data widely used in LLM training; an adjacent peer in ethical/open data supply, though focused on web text rather than community datasets.
- Surge AI: Surge AI operates an AI data labeling and dataset platform focused on RLHF and frontier model evaluations; competes for premium AI data buyers with a contributor-friendly positioning.
- Snorkel AI: Snorkel AI provides data-centric AI development tooling including programmatic labeling and dataset management; serves a similar AI builder audience with adjacent data platform capabilities.
Others
- Open Data Institute: The Open Data Institute advocates for and curates open, ethical data ecosystems; thematically aligned with MDC's mission and a potential policy/distribution partner in the trusted data commons space.
Market position
Strengths4 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights6 records
Customer concentration
Mozilla Data Collective social profiles
Digital presenceMozilla Data Collective compliance and trust
Trust signalCompliance2 records
Mozilla Data Collective financial estimates
Financial estimateRevenue estimate
Valuation estimate
Mozilla Data Collective leadership team
Management profileNumber of profiles
Mozilla Data Collective funding detail
Funding detailFunding overview
Funding rounds1 record
Investors1 record
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Mozilla Data Collective M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Mozilla Data Collective
What does Mozilla Data Collective do?
Mozilla Data Collective operates a cloud-based data sharing platform and marketplace for AI training datasets, enabling data contributors to upload, license, and distribute datasets under open, community-governed, or compensated models. Data consumers—including AI/ML researchers, developers, and organizations—discover and download multilingual, multicultural, and multimodal datasets through a web interface, REST API, and Python SDK. The platform charges a 5% fee on compensated dataset transactions, with 100% of dataset license fees going to contributors.
Is Mozilla Data Collective a public or private company?
Mozilla Data Collective is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was Mozilla Data Collective founded?
Mozilla Data Collective was founded in 2025.
Where is Mozilla Data Collective based?
Mozilla Data Collective is headquartered in London, United Kingdom, in the Europe region.
How does Mozilla Data Collective make money?
Two revenue lines are on record. Platform Fee on Dataset Transactions are the primary driver. The others are dataset Licensing Fees.
Who are Mozilla Data Collective's main competitors?
Direct peers on record are Hugging Face, Kaggle and Roboflow. Broad incumbents are AWS Data Exchange, Scale AI and Appen. Emerging players are Common Crawl, Surge AI and Snorkel AI. Open Data Institute is listed as an others.
Does Mozilla Data Collective have an API?
Yes. The Mozilla Data Collective REST API provides programmatic access to datasets. Base URL: https://dev.mozilladatacollective.com/api. Authentication uses Bearer tokens in the Authorization header. Key endpoints include GET /datasets/:datasetId (retrieve dataset details) and POST /datasets/:datasetId/download (create a download session returning a presigned URL). Users must agree to dataset terms through the web interface before downloading via API. Rate limiting is at the organization level: 30 dataset downloads per day per organization, with a 12-hour expiration on presigned URLs. A Python SDK is available for easier integration. Direct storage downloads from S3/R2 are supported with resumable (range request) capability. Developer documentation is at dev.mozilladatacollective.com/api-reference/docs.
What industry is Mozilla Data Collective in?
Mozilla Data Collective's product category is AI Data Marketplace. Its primary akta.pro industry code is HDAAALAE, Dataset Discovery, Marketplaces & Licensing, with a secondary code of HDAEABAH, Data Sharing, Data Exchange & Data Marketplace Platforms. Its NAICS code is 51821 and its SIC code is 7370.