ELDA - Evaluations and Language resources Distribution Agency
ELDA is the Paris-based operational arm of ELRA, distributing and licensing validated language resources (speech corpora, lexicons, terminological data) across 270+ languages to researchers, language technology vendors, publishers, and EU institutions.
- Company typePrivate
- Founded1995
- HeadquartersParis, France
- Headcount51–100
- GTM typeB2B
- OfferingDigital Commerce or Content
What ELDA - Evaluations and Language resources Distribution Agency does
ELDA (Evaluations and Language Resources Distribution Agency) is a Paris-based private SME established in February 1995 as the operational arm of the European Language Resources Association (ELRA). It operates as the central European distributor and licensor of validated language data — speech corpora, written corpora, lexicons, terminological resources, and multilingual/multimodal datasets across more than 270 languages. Its primary product is the ELRA Catalogue (catalog.elra.info), supplemented by an R&D Catalogue (academic pricing) and a members-only Universal Catalogue. Alongside distribution, ELDA runs the ISLRN persistent identifier system (jointly with LDC), which has assigned 3,353 identifiers since 2014 and underpins citable attribution across the European NLP research community.
The company monetizes through a mix of one-time/perpetual catalogue licenses (with member, non-member, academic, and commercial pricing tiers), annual ELRA membership subscriptions, 3-month evaluation licenses, and a small share of free academic resources. Around this core distribution model, ELDA builds a lifecycle of services: identification, production, validation, distribution, legal support, license tooling (License Wizard), GDPR-compliant de-identification/anonymisation, and Data Management Plan support. Non-commercial revenue is anchored by multi-year European Commission project participation, notably the 3-year Common European Language Data Space (LDS) launched January 2023 alongside DFKI, ILSP, and SIA Tilde; the European Multilingual Web (EMW) under DIGITAL Europe; and the Language Technologies Solutions ASR project led by Brno University of Technology.
Go-to-market is hybrid: catalog-elra.info plus signed order forms (notably still fax/mail-based for ordering), email newsletter, Twitter, and an event-led presence through the biennial LREC conference series (LREC-COLING 2024 scheduled in Turin). Channel partnerships extend the catalogue through Datatang (67 Asian/Middle Eastern speech resources) and Lexicala (50-language multilingual lexical data). Target customers are academic and research institutions (primary), European language technology vendors, publishers/language service providers, and EU/public-sector multilingual service buyers. ELDA is privately held with no disclosed external funding, employs 86 people, and is led by CEO Khalid Choukri.
ELDA - Evaluations and Language resources Distribution Agency firmographics
Firmographics- Name
- ELDA - Evaluations and Language resources Distribution Agency
- Legal name
- ELDA - Evaluations and Language Resources Distribution Agency
- Website
- https://elda.org
- Company type
- Private
- Founded year
- 1995
- Operating status
- Operating
- Headcount range
- 51–100 employees
- Short description
- ELDA is the Paris-based operational arm of ELRA, distributing and licensing validated language resources (speech corpora, lexicons, terminological data) across 270+ languages to researchers, language technology vendors, publishers, and EU institutions.
- Ownership category
- akta.pro rank
Where ELDA - Evaluations and Language resources Distribution Agency is headquartered
LocationHeadquarters
- HQ city
- Paris
- HQ country
- France
- HQ region
- Europe
Offices1 record
Markets served
ELDA - Evaluations and Language resources Distribution Agency business model
Business model- GTM type
- B2B
- Offering type
- Digital Commerce or Content
- Cost components
- Personnel, Operations, Technology or R&D, Marketing or Sales, Infrastructure, Others
Revenue model
- Language Resources Sales: Sale of language resources (speech corpora, written corpora, lexicons, terminological resources) through the ELRA Catalogue. Pricing varies by resource type, organization type (academic/commercial), and membership status.
- Membership Fees: Annual membership fees for ELRA that provide access to discounts on language resources and the Universal Catalogue.
- Licensing Services: Licensing of language resources for research, commercial, or evaluation use (3-month evaluation period available).
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Subscription | Annual | Member pricing for academic/research use |
| One time/ perpetual license | Pay-as-you-go | Non-member academic pricing |
| One time/ perpetual license | Pay-as-you-go | Commercial use pricing |
| Transaction based/ take rate | Pay-as-you-go | Evaluation license |
| Freemium | Pay-as-you-go | Free resources for academic research |
Go-to-market motion2 records
Distribution channels4 records
Marketing channels6 records
ELDA - Evaluations and Language resources Distribution Agency product offering
Product offeringCore offering
ELDA distributes validated language resources (speech corpora, written corpora, lexicons, and terminological resources) from third-party producers through the ELRA Catalogue at catalog.elra.info, selling both perpetual and evaluation licenses to academic and commercial buyers. Complementing the catalogue, ELDA delivers a service portfolio covering the full language-resource lifecycle: Identification of existing resources, Production of new corpora, Validation against standards, and Distribution with licensing/Legal Helpdesk support.
Product overview
ELDA (Evaluations and Language resources Distribution Agency) is ELRA's operational body, established in 1995 as an SME in Paris, specializing in Human Language Technologies. The core offering is the ELRA Catalogue of Language Resources - a comprehensive marketplace for speech corpora, written corpora, lexicons, and terminological resources distributed globally. Supporting this are three interconnected catalogues: the main ELRA Catalogue for commercial and research purchases, the R&D Catalogue for affordable academic use, and the Universal Catalogue for member-only resource discovery. ISLRN provides unique persistent identifiers for proper resource citation. The service portfolio spans the full language resource lifecycle: Identification (discovering resources), Production (creating new resources), Validation (ensuring quality/standards), and Distribution (licensing and delivery). Key projects include the Common European Language Data Space (LDS) platform for multilingual data sharing, the European Multilingual Web (EMW) for automated translation, and Language Technologies Solutions for speech recognition. Additional services include the License Wizard, Legal Support Helpdesk, De-identification/Anonymisation, and HLT Evaluation services. ELDA also publishes the Language Resources and Evaluation Journal and organizes the LREC conference series.
Differentiator
Problem solved
Functional benefit
Brands
- ELRA Catalogue of Language Resources: ELRA's Language Resources Catalogue containing speech, multimodal, written lexica/corpus, and terminological resources available for purchase
- ISLRN
- R&D Catalogue
- Universal Catalogue
- LREC-COLING
- Common European Language Data Space (LDS)
Products and services
- ELRA Catalogue of Language Resources Comprehensive online catalogue at catalog.elra.info distributing speech corpora, multimodal corpora, written corpora, lexicons, and terminological resources on a pay-per-resource perpetual licence basis to academic and commercial buyers worldwide.
- R&D Catalogue Academic-research subset of the ELRA Catalogue offering over 200 language resources at 500 EUR and below, targeted at universities and public research laboratories.
- Universal Catalogue ELRA member-only service providing discovery and search-aid information about identified language resources awaiting full catalogue inclusion.
- ISLRN (International Standard Language Resource Number) Unique persistent identifier system for language resources, jointly administered with the Lingu Data Consortium, ensuring correct identification and citation of resources in research papers, products, and evaluation benchmarks.
- Identification Services Professional service identifying existing language resources not yet in the catalogue and negotiating with rights holders to include them in ELDA's distribution pipeline.
- Production Services Professional service producing new language resources for natural language processing, speech systems, software localisation, language services, and electronic publishing.
- Validation Services Standards-compliance and best-practices validation of language resources, including review by ELDA's Validation Committee, ensuring data quality and interoperability.
- Distribution Services Licensing and distribution service covering pricing strategy, licence agreement generation, and delivery of language resources to end users globally.
- ELRA License Wizard Online tool that helps users determine and generate appropriate licence agreements (research, commercial, evaluation) for language resources purchased via ELDA.
- Legal Support Helpdesk Legal-assistance service for language resource procurement, including contract negotiation with providers, licensing guidance, and GDPR/data protection compliance.
- De-identification and Anonymisation Services Personal-data protection and anonymisation service for language resources, supporting GDPR and EUDPR compliance for research and commercial users.
- Data Management Plan (DMP) Assistance with data-management planning for language-resource projects, supporting compliance with funder and institutional DMP requirements.
- Common European Language Data Space (LDS) Platform Three-year EU-funded platform and marketplace for collection, creation, sharing and re-use of multilingual and multimodal language data across Europe, co-developed by ELDA with DFKI, Athena RIC, and SIA Tilde.
- Bitext Lexical Datasets and Bitext Synthetic Data 66 monolingual lexical datasets and synthetic data products added to the catalogue in July 2023, covering multiple verticals (20 verticals) in English and Spanish for intent detection and NLP applications.
Quantifiable outcome
- 3,353 ISLRN numbers assigned to language resources worldwide
- +2 more outcomes
Companies that use ELDA - Evaluations and Language resources Distribution Agency
Customer profileNamed customers3 records
Segments4 records
Ideal customer profiles4 records
ELDA - Evaluations and Language resources Distribution Agency technology and API
TechnologyTechnology focussed No
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability9 records
Feature3 records
ELDA - Evaluations and Language resources Distribution Agency partnerships and signals
Strategic signalPartnerships
15 partnerships are on record, tiered major, core and minor.
- Lexicala by K DictionariesmajorDistribution agreement for Lexicala's multilingual lexical data covering 50 languages. ELDA distributes Lexicala's high-quality lexical data for language learning, machine translation, and NLP applications.
- German Research Center for Artificial Intelligence (DFKI)coreDFKI is the coordinator of the Common European Language Data Space (LDS) project, a 3-year European Commission initiative. ELDA is one of four consortium partners working on establishing a European platform and marketplace for multilingual and multimodal language data.
- Athena Research and Innovation Center (ILSP)coreILSP is a consortium partner in the LDS project, responsible for developing and implementing a sustainable language data ecosystem blueprint and Language Data Space deployment.
- SIA TildecoreSIA Tilde participates in the LDS project, focusing on implementation of multi-stakeholder data and services governance scheme and proof-of-deployment-concept projects.
- Brno University of Technology (BUT)majorCoordinator of Language Technologies Solutions project. ELDA leads market study on Automatic Speaker Recognition solutions and coordinates collection of speech data for under-resourced European languages.
- Tilde (Latvia)majorCoordinator of European Multilingual Web (EMW) project within Digital Europe Programme. ELDA responsible for helpdesk management supporting automated website translation solutions.
- DatatangmajorLanguage resources distribution agreement for 67 speech resources designed to boost speech recognition. ELDA strengthens its position as worldwide distribution center while Datatang gains European market visibility.
- Linguistic Data Consortium (LDC)coreLong-standing partnership with LDC for distribution of language resources, joint ISLRN implementation, and collaborative projects including NETDC production/distribution and CoNLL joint distribution.
- CLARIN ERICcoreCollaboration agreement between ELRA and CLARIN ERIC for joint activities in language resources infrastructure and sharing.
- INSA Rouen NormandieminorPartnership to release the Annotated tweet corpus in Arabizi, French and English for research purposes.
- FID LinguistikminorPartnership to grant licenses for written corpora to German researchers, facilitating access to language resources for the German research community.
- IULA (Universitat Pompeu Fabra)minorIULA has adopted ISLRN for identification of language resources.
- Resource Management Agency (RMA)minorRMA has adopted ISLRN for identification of language resources.
- Vigdís International Centre of MultilingualismminorPartnership for multilingualism and intercultural understanding initiatives.
- ILC-ELRAminorPartnership with ILC for language resources activities.
Scale indicators6 records
Recent moves6 records
Expansion highlights5 records
ELDA - Evaluations and Language resources Distribution Agency competitors and assessment
Company assessmentBroad incumbents
- Defined.ai: Defined.ai (formerly DefinedCrowd) is a commercial marketplace for training data covering speech, text, and image. It competes with ELDA for the speech and NLP training data budgets of large enterprise and AI-lab buyers.
- Appen: Appen is a publicly traded data annotation and dataset provider for AI/ML, covering text, speech, and image. While ELDA focuses on curated, validated research-grade corpora, Appen competes for the same enterprise speech/text data budgets at a much larger scale.
Direct peers
- CLARIN ERIC: CLARIN is the European research infrastructure for language resources, with a federated catalogue of corpora and tools across European universities. ELDA has a formal collaboration agreement with CLARIN and they overlap significantly in serving academic NLP researchers with discoverable, licensed language data.
- META-SHARE Network: META-SHARE is the federated network of European language resource repositories in which ELDA participates. It is a parallel infrastructure for sharing language data and overlaps heavily with ELDA's catalogue in target audience and resource types.
- European Language Grid: European Language Grid is an EU-funded platform cataloguing language technology assets and resources across Europe. It is a closely adjacent infrastructure to ELDA's catalogue and serves the same multilingual European market.
- Tilde: Tilde is a Latvian language technology company and consortium partner with ELDA on multiple EU projects (EMW, LDS). Both operate in European multilingual NLP, deliver language data services, and target enterprise/government buyers for translation and speech tools.
- Linguistic Data Consortium (LDC): LDC at the University of Pennsylvania is the closest peer: a non-profit catalogue and distribution body for speech corpora, text corpora, and lexicons. ELDA and LDC are explicit co-administrators of ISLRN and joint distributors, and they serve the same academic and commercial NLP buyers.
Emerging players
- Hugging Face Datasets: Hugging Face hosts a large open catalogue of NLP and speech datasets and has become the default distribution channel for ML practitioners. It overlaps with ELDA's buyer base (NLP researchers and developers) but operates on an open-source, instant-download model rather than licensed catalogue sales.
- Datatang: Datatang is a Chinese AI data provider whose 67-resource speech catalogue is distributed by ELDA in Europe. It is comparable as a seller of speech training data, although geographic and buyer profiles differ.
- K Dictionaries (Lexicala): Lexicala by K Dictionaries provides multilingual lexical data across 50 languages and distributes through ELDA. It is a focused peer in the lexicons/terminology sub-segment of ELDA's catalogue.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights7 records
Customer concentration
ELDA - Evaluations and Language resources Distribution Agency social profiles
Digital presenceELDA - Evaluations and Language resources Distribution Agency compliance and trust
Trust signalCompliance2 records
ELDA - Evaluations and Language resources Distribution Agency financial estimates
Financial estimateRevenue estimate
Valuation estimate
ELDA - Evaluations and Language resources Distribution Agency leadership team
Management profileNumber of profiles
Profiles1 record
ELDA - Evaluations and Language resources Distribution Agency funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
ELDA - Evaluations and Language resources Distribution Agency M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about ELDA - Evaluations and Language resources Distribution Agency
What does ELDA - Evaluations and Language resources Distribution Agency do?
ELDA distributes validated language resources (speech corpora, written corpora, lexicons, and terminological resources) from third-party producers through the ELRA Catalogue at catalog.elra.info, selling both perpetual and evaluation licenses to academic and commercial buyers. Complementing the catalogue, ELDA delivers a service portfolio covering the full language-resource lifecycle: Identification of existing resources, Production of new corpora, Validation against standards, and Distribution with licensing/Legal Helpdesk support.
Is ELDA - Evaluations and Language resources Distribution Agency a public or private company?
ELDA - Evaluations and Language resources Distribution Agency is a private company. It is classified as management employee owned and is currently operating.
When was ELDA - Evaluations and Language resources Distribution Agency founded?
ELDA - Evaluations and Language resources Distribution Agency was founded in 1995. It employs 51 to 100 people.
Where is ELDA - Evaluations and Language resources Distribution Agency based?
ELDA - Evaluations and Language resources Distribution Agency is headquartered in Paris, France, in the Europe region.
How does ELDA - Evaluations and Language resources Distribution Agency make money?
Three revenue lines are on record. Language Resources Sales are the primary driver. The others are membership Fees and licensing Services.
Who are ELDA - Evaluations and Language resources Distribution Agency's main competitors?
Broad incumbents on record are Defined.ai and Appen. Direct peers are CLARIN ERIC, META-SHARE Network, European Language Grid, Tilde and Linguistic Data Consortium (LDC). Emerging players are Hugging Face Datasets, Datatang and K Dictionaries (Lexicala).
Does ELDA - Evaluations and Language resources Distribution Agency have an API?
No public API is recorded for ELDA - Evaluations and Language resources Distribution Agency.