Developer docs
API playgroundTry for free, no card

Search company profiles

ELDA - Evaluations and Language resources Distribution Agency

Full company profile

uuid004c2yg

Namestring
ELDA - Evaluations and Language resources Distribution Agency
Legal namestring
ELDA - Evaluations and Language Resources Distribution Agency
Websiteurl
elda.org
Company typeenum
Private
Founded yearint
1995
Descriptiontext

ELDA (Evaluations and Language Resources Distribution Agency) is a Paris-based private SME established in February 1995 as the operational arm of the European Language Resources Association (ELRA). It operates as the central European distributor and licensor of validated language data — speech corpora, written corpora, lexicons, terminological resources, and multilingual/multimodal datasets across more than 270 languages. Its primary product is the ELRA Catalogue (catalog.elra.info), supplemented by an R&D Catalogue (academic pricing) and a members-only Universal Catalogue. Alongside distribution, ELDA runs the ISLRN persistent identifier system (jointly with LDC), which has assigned 3,353 identifiers since 2014 and underpins citable attribution across the European NLP research community.

The company monetizes through a mix of one-time/perpetual catalogue licenses (with member, non-member, academic, and commercial pricing tiers), annual ELRA membership subscriptions, 3-month evaluation licenses, and a small share of free academic resources. Around this core distribution model, ELDA builds a lifecycle of services: identification, production, validation, distribution, legal support, license tooling (License Wizard), GDPR-compliant de-identification/anonymisation, and Data Management Plan support. Non-commercial revenue is anchored by multi-year European Commission project participation, notably the 3-year Common European Language Data Space (LDS) launched January 2023 alongside DFKI, ILSP, and SIA Tilde; the European Multilingual Web (EMW) under DIGITAL Europe; and the Language Technologies Solutions ASR project led by Brno University of Technology.

Go-to-market is hybrid: catalog-elra.info plus signed order forms (notably still fax/mail-based for ordering), email newsletter, Twitter, and an event-led presence through the biennial LREC conference series (LREC-COLING 2024 scheduled in Turin). Channel partnerships extend the catalogue through Datatang (67 Asian/Middle Eastern speech resources) and Lexicala (50-language multilingual lexical data). Target customers are academic and research institutions (primary), European language technology vendors, publishers/language service providers, and EU/public-sector multilingual service buyers. ELDA is privately held with no disclosed external funding, employs 86 people, and is led by CEO Khalid Choukri.

Short descriptiontext

ELDA is the Paris-based operational arm of ELRA, distributing and licensing validated language resources (speech corpora, lexicons, terminological data) across 270+ languages to researchers, language technology vendors, publishers, and EU institutions.

Operating statusenum
Operating
Ownership categoryenum
Headcount rangeband
51–100
akta.pro rankint
HeadquartersParis, France
HQ citystring
Paris
HQ countrystring
France
HQ regionstring
Europe
Markets served

Serves global market

Offices1 record

Each record includes

City, Country, Type, Description, Source

Keyword5 values
language resource distribution, speech corpus licensing, computational linguistics data, multilingual data marketplace, NLP data catalogue
NAICS code1 code
  • Newspaper, Periodical, Book, and Directory Publishers5131
SIC code1 code
  • Miscellaneous Publishing2741
Product category
Language Resources Distribution
Social media profiles2 records
GTM motion2 records

Each record includes

Type, Description, Source

Revenue model3 records
1Language Resources Sales
TypeOne Time License
Description

Sale of language resources (speech corpora, written corpora, lexicons, terminological resources) through the ELRA Catalogue. Pricing varies by resource type, organization type (academic/commercial), and membership status.

elda.org
2Membership Fees
TypeSubscription Recurring
Description

Annual membership fees for ELRA that provide access to discounts on language resources and the Universal Catalogue.

elda.org
3Licensing Services
TypeLicensing Royalties
Description

Licensing of language resources for research, commercial, or evaluation use (3-month evaluation period available).

elda.org
Marketing channels6 records

Each record includes

Title, Type, Stage, Description, Source

Distribution channels4 records

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components6 values
Personnel, Operations, Technology or R&D, Marketing or Sales, Infrastructure, Others
Pricing details5 tiers
1Member pricing for academic/research use
ModelSubscriptionBilling cadenceAnnual
Notes

Discounted rates for ELRA members for research purposes

elda.org
2Non-member academic pricing
ModelOne time/ perpetual licenseBilling cadencePay-as-you-go
Notes

Full public pricing for academic institutions not holding ELRA membership

elda.org
3Commercial use pricing
ModelOne time/ perpetual licenseBilling cadencePay-as-you-go
Notes

Higher pricing tier for commercial/organization use

elda.org
4Evaluation license
ModelTransaction based/ take rateBilling cadencePay-as-you-go
Notes

3-month evaluation period available for language resources

elda.org
5Free resources for academic research
ModelFreemiumBilling cadencePay-as-you-go
Notes

MEDIA data available for free for academic research. Some resources available free except shipment.

elda.org
GTM typeB2B
B2B
Offering typeDigital Commerce or Conte…
Digital Commerce or Content
Brand1 of 6 records shown
1ELRA Catalogue of Language Resources
Description

ELRA's Language Resources Catalogue containing speech, multimodal, written lexica/corpus, and terminological resources available for purchase

elda.org
+5 more records
Core offering1 text field

ELDA distributes validated language resources (speech corpora, written corpora, lexicons, and terminological resources) from third-party producers through the ELRA Catalogue at catalog.elra.info, selling both perpetual and evaluation licenses to academic and commercial buyers. Complementing the catalogue, ELDA delivers a service portfolio covering the full language-resource lifecycle: Identification of existing resources, Production of new corpora, Validation against standards, and Distribution with licensing/Legal Helpdesk support.

Differentiator
Functional benefit
Problem solved
Quantifiable outcome1 of 3 values shown
  • 3,353 ISLRN numbers assigned to language resources worldwide
+2 more records
Product overview1 text field

ELDA (Evaluations and Language resources Distribution Agency) is ELRA's operational body, established in 1995 as an SME in Paris, specializing in Human Language Technologies. The core offering is the ELRA Catalogue of Language Resources - a comprehensive marketplace for speech corpora, written corpora, lexicons, and terminological resources distributed globally. Supporting this are three interconnected catalogues: the main ELRA Catalogue for commercial and research purchases, the R&D Catalogue for affordable academic use, and the Universal Catalogue for member-only resource discovery. ISLRN provides unique persistent identifiers for proper resource citation. The service portfolio spans the full language resource lifecycle: Identification (discovering resources), Production (creating new resources), Validation (ensuring quality/standards), and Distribution (licensing and delivery). Key projects include the Common European Language Data Space (LDS) platform for multilingual data sharing, the European Multilingual Web (EMW) for automated translation, and Language Technologies Solutions for speech recognition. Additional services include the License Wizard, Legal Support Helpdesk, De-identification/Anonymisation, and HLT Evaluation services. ELDA also publishes the Language Resources and Evaluation Journal and organizes the LREC conference series.

Product and service14 records
1ELRA Catalogue of Language Resources
CategoryLanguage Resources Distribution
Description

Comprehensive online catalogue at catalog.elra.info distributing speech corpora, multimodal corpora, written corpora, lexicons, and terminological resources on a pay-per-resource perpetual licence basis to academic and commercial buyers worldwide.

2R&D Catalogue
CategoryLanguage Resources Distribution
Description

Academic-research subset of the ELRA Catalogue offering over 200 language resources at 500 EUR and below, targeted at universities and public research laboratories.

3Universal Catalogue
CategoryLanguage Resources Distribution
Description

ELRA member-only service providing discovery and search-aid information about identified language resources awaiting full catalogue inclusion.

4ISLRN (International Standard Language Resource Number)
CategoryStandards and Identification
Description

Unique persistent identifier system for language resources, jointly administered with the Lingu Data Consortium, ensuring correct identification and citation of resources in research papers, products, and evaluation benchmarks.

5Identification Services
CategoryLanguage Resource Services
Description

Professional service identifying existing language resources not yet in the catalogue and negotiating with rights holders to include them in ELDA's distribution pipeline.

6Production Services
CategoryLanguage Resource Services
Description

Professional service producing new language resources for natural language processing, speech systems, software localisation, language services, and electronic publishing.

7Validation Services
CategoryLanguage Resource Services
Description

Standards-compliance and best-practices validation of language resources, including review by ELDA's Validation Committee, ensuring data quality and interoperability.

8Distribution Services
CategoryLanguage Resource Services
Description

Licensing and distribution service covering pricing strategy, licence agreement generation, and delivery of language resources to end users globally.

9ELRA License Wizard
CategoryLicensing Tools
Description

Online tool that helps users determine and generate appropriate licence agreements (research, commercial, evaluation) for language resources purchased via ELDA.

10Legal Support Helpdesk
CategoryLegal and Compliance Services
Description

Legal-assistance service for language resource procurement, including contract negotiation with providers, licensing guidance, and GDPR/data protection compliance.

11De-identification and Anonymisation Services
CategoryLegal and Compliance Services
Description

Personal-data protection and anonymisation service for language resources, supporting GDPR and EUDPR compliance for research and commercial users.

12Data Management Plan (DMP)
CategoryLanguage Resource Services
Description

Assistance with data-management planning for language-resource projects, supporting compliance with funder and institutional DMP requirements.

13Common European Language Data Space (LDS) Platform
CategoryEuropean Data Platform
Description

Three-year EU-funded platform and marketplace for collection, creation, sharing and re-use of multilingual and multimodal language data across Europe, co-developed by ELDA with DFKI, Athena RIC, and SIA Tilde.

14Bitext Lexical Datasets and Bitext Synthetic Data
CategoryLexical Resources
Description

66 monolingual lexical datasets and synthetic data products added to the catalogue in July 2023, covering multiple verticals (20 verticals) in English and Spanish for intent detection and NLP applications.

Scale indicator6 records

Each record includes

Type, Value, Description, Source

Partnership15 partners
Strategic tierMajorTypeChannel Partner/ Reseller/ DistributorAnnounced on2023-10-12
Description

Distribution agreement for Lexicala's multilingual lexical data covering 50 languages. ELDA distributes Lexicala's high-quality lexical data for language learning, machine translation, and NLP applications.

2German Research Center for Artificial Intelligence (DFKI)
Strategic tierCoreTypeStrategic or Co-development PartnerAnnounced on2023-01-19
Description

DFKI is the coordinator of the Common European Language Data Space (LDS) project, a 3-year European Commission initiative. ELDA is one of four consortium partners working on establishing a European platform and marketplace for multilingual and multimodal language data.

elda.org
Strategic tierCoreTypeStrategic or Co-development PartnerAnnounced on2023-01-19
Description

ILSP is a consortium partner in the LDS project, responsible for developing and implementing a sustainable language data ecosystem blueprint and Language Data Space deployment.

Strategic tierCoreTypeStrategic or Co-development PartnerAnnounced on2023-01-19
Description

SIA Tilde participates in the LDS project, focusing on implementation of multi-stakeholder data and services governance scheme and proof-of-deployment-concept projects.

Strategic tierMajorTypeStrategic or Co-development PartnerAnnounced on2022-12-13
Description

Coordinator of Language Technologies Solutions project. ELDA leads market study on Automatic Speaker Recognition solutions and coordinates collection of speech data for under-resourced European languages.

Strategic tierMajorTypeStrategic or Co-development PartnerAnnounced on2022-12-12
Description

Coordinator of European Multilingual Web (EMW) project within Digital Europe Programme. ELDA responsible for helpdesk management supporting automated website translation solutions.

Strategic tierMajorTypeChannel Partner/ Reseller/ DistributorAnnounced on2022-10-27
Description

Language resources distribution agreement for 67 speech resources designed to boost speech recognition. ELDA strengthens its position as worldwide distribution center while Datatang gains European market visibility.

Strategic tierCoreTypeStrategic or Co-development Partner
Description

Long-standing partnership with LDC for distribution of language resources, joint ISLRN implementation, and collaborative projects including NETDC production/distribution and CoNLL joint distribution.

9CLARIN ERIC
Strategic tierCoreTypeStrategic or Co-development Partner
Description

Collaboration agreement between ELRA and CLARIN ERIC for joint activities in language resources infrastructure and sharing.

elda.org
Strategic tierMinorTypeStrategic or Co-development Partner
Description

Partnership to release the Annotated tweet corpus in Arabizi, French and English for research purposes.

11FID Linguistik
Strategic tierMinorTypeChannel Partner/ Reseller/ Distributor
Description

Partnership to grant licenses for written corpora to German researchers, facilitating access to language resources for the German research community.

elda.org
12IULA (Universitat Pompeu Fabra)
Strategic tierMinorTypeStrategic or Co-development Partner
Description

IULA has adopted ISLRN for identification of language resources.

elda.org
Strategic tierMinorTypeStrategic or Co-development Partner
Description

RMA has adopted ISLRN for identification of language resources.

14Vigdís International Centre of Multilingualism
Strategic tierMinorTypeStrategic or Co-development Partner
Description

Partnership for multilingualism and intercultural understanding initiatives.

elda.org
Strategic tierMinorTypeStrategic or Co-development Partner
Description

Partnership with ILC for language resources activities.

Recent move6 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight5 records

Each record includes

Type, Description

Peers10 records
TypeBroad incumbent
Description

Defined.ai (formerly DefinedCrowd) is a commercial marketplace for training data covering speech, text, and image. It competes with ELDA for the speech and NLP training data budgets of large enterprise and AI-lab buyers.

2CLARIN ERIC
TypeDirect peer
Description

CLARIN is the European research infrastructure for language resources, with a federated catalogue of corpora and tools across European universities. ELDA has a formal collaboration agreement with CLARIN and they overlap significantly in serving academic NLP researchers with discoverable, licensed language data.

TypeEmerging player
Description

Hugging Face hosts a large open catalogue of NLP and speech datasets and has become the default distribution channel for ML practitioners. It overlaps with ELDA's buyer base (NLP researchers and developers) but operates on an open-source, instant-download model rather than licensed catalogue sales.

4META-SHARE Network
TypeDirect peer
Description

META-SHARE is the federated network of European language resource repositories in which ELDA participates. It is a parallel infrastructure for sharing language data and overlaps heavily with ELDA's catalogue in target audience and resource types.

5European Language Grid
TypeDirect peer
Description

European Language Grid is an EU-funded platform cataloguing language technology assets and resources across Europe. It is a closely adjacent infrastructure to ELDA's catalogue and serves the same multilingual European market.

TypeDirect peer
Description

Tilde is a Latvian language technology company and consortium partner with ELDA on multiple EU projects (EMW, LDS). Both operate in European multilingual NLP, deliver language data services, and target enterprise/government buyers for translation and speech tools.

TypeDirect peer
Description

LDC at the University of Pennsylvania is the closest peer: a non-profit catalogue and distribution body for speech corpora, text corpora, and lexicons. ELDA and LDC are explicit co-administrators of ISLRN and joint distributors, and they serve the same academic and commercial NLP buyers.

TypeBroad incumbent
Description

Appen is a publicly traded data annotation and dataset provider for AI/ML, covering text, speech, and image. While ELDA focuses on curated, validated research-grade corpora, Appen competes for the same enterprise speech/text data budgets at a much larger scale.

TypeEmerging player
Description

Datatang is a Chinese AI data provider whose 67-resource speech catalogue is distributed by ELDA in Europe. It is comparable as a seller of speech training data, although geographic and buyer profiles differ.

TypeEmerging player
Description

Lexicala by K Dictionaries provides multilingual lexical data across 50 languages and distributes through ELDA. It is a focused peer in the lexicons/terminology sub-segment of ELDA's catalogue.

Market position
Strengths5 records

Each record includes

Headline, Details, Source

Weaknesses5 records

Each record includes

Headline, Details, Source

Competitive moat5 records

Each record includes

Type, Details

Key risks6 records

Each record includes

Headline, Details, Source

Key highlights7 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Named customers3 records

Each record includes

Name, Industry, Type, Use case, Source, UUID

Segment4 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile4 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
No
API detail
Has APIbool
No

Docs URL, Description

AI capability9 records

Each record includes

Type, Description, Source

AI maturity
App detail

Has app

Feature3 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
Profiles1 record

Each record includes

Name, Designation, Designation category, Overview, Profile commentary, Source

No data
Compliance2 records

Each record includes

Name, Class, Description

Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

ELDA - Evaluations and Language resources Distribution Agency

Language Resources Distributionelda.org

ELDA is the Paris-based operational arm of ELRA, distributing and licensing validated language resources (speech corpora, lexicons, terminological data) across 270+ languages to researchers, language technology vendors, publishers, and EU institutions.

What ELDA - Evaluations and Language resources Distribution Agency does

ELDA (Evaluations and Language Resources Distribution Agency) is a Paris-based private SME established in February 1995 as the operational arm of the European Language Resources Association (ELRA). It operates as the central European distributor and licensor of validated language data — speech corpora, written corpora, lexicons, terminological resources, and multilingual/multimodal datasets across more than 270 languages. Its primary product is the ELRA Catalogue (catalog.elra.info), supplemented by an R&D Catalogue (academic pricing) and a members-only Universal Catalogue. Alongside distribution, ELDA runs the ISLRN persistent identifier system (jointly with LDC), which has assigned 3,353 identifiers since 2014 and underpins citable attribution across the European NLP research community.

The company monetizes through a mix of one-time/perpetual catalogue licenses (with member, non-member, academic, and commercial pricing tiers), annual ELRA membership subscriptions, 3-month evaluation licenses, and a small share of free academic resources. Around this core distribution model, ELDA builds a lifecycle of services: identification, production, validation, distribution, legal support, license tooling (License Wizard), GDPR-compliant de-identification/anonymisation, and Data Management Plan support. Non-commercial revenue is anchored by multi-year European Commission project participation, notably the 3-year Common European Language Data Space (LDS) launched January 2023 alongside DFKI, ILSP, and SIA Tilde; the European Multilingual Web (EMW) under DIGITAL Europe; and the Language Technologies Solutions ASR project led by Brno University of Technology.

Go-to-market is hybrid: catalog-elra.info plus signed order forms (notably still fax/mail-based for ordering), email newsletter, Twitter, and an event-led presence through the biennial LREC conference series (LREC-COLING 2024 scheduled in Turin). Channel partnerships extend the catalogue through Datatang (67 Asian/Middle Eastern speech resources) and Lexicala (50-language multilingual lexical data). Target customers are academic and research institutions (primary), European language technology vendors, publishers/language service providers, and EU/public-sector multilingual service buyers. ELDA is privately held with no disclosed external funding, employs 86 people, and is led by CEO Khalid Choukri.

ELDA - Evaluations and Language resources Distribution Agency firmographics

Firmographics
Name
ELDA - Evaluations and Language resources Distribution Agency
Legal name
ELDA - Evaluations and Language Resources Distribution Agency
Website
https://elda.org
Company type
Private
Founded year
1995
Operating status
Operating
Headcount range
51–100 employees
Short description
ELDA is the Paris-based operational arm of ELRA, distributing and licensing validated language resources (speech corpora, lexicons, terminological data) across 270+ languages to researchers, language technology vendors, publishers, and EU institutions.
Ownership category
akta.pro rank

Where ELDA - Evaluations and Language resources Distribution Agency is headquartered

Location

Headquarters

HQ city
Paris
HQ country
France
HQ region
Europe

Offices1 record

Markets served

ELDA - Evaluations and Language resources Distribution Agency business model

Business model
GTM type
B2B
Offering type
Digital Commerce or Content
Cost components
Personnel, Operations, Technology or R&D, Marketing or Sales, Infrastructure, Others

Revenue model

  1. Language Resources Sales: Sale of language resources (speech corpora, written corpora, lexicons, terminological resources) through the ELRA Catalogue. Pricing varies by resource type, organization type (academic/commercial), and membership status.
  2. Membership Fees: Annual membership fees for ELRA that provide access to discounts on language resources and the Universal Catalogue.
  3. Licensing Services: Licensing of language resources for research, commercial, or evaluation use (3-month evaluation period available).

Pricing tiers

ModelBillingPrice
SubscriptionAnnualMember pricing for academic/research use
One time/ perpetual licensePay-as-you-goNon-member academic pricing
One time/ perpetual licensePay-as-you-goCommercial use pricing
Transaction based/ take ratePay-as-you-goEvaluation license
FreemiumPay-as-you-goFree resources for academic research

Go-to-market motion2 records

Distribution channels4 records

Marketing channels6 records

ELDA - Evaluations and Language resources Distribution Agency product offering

Product offering

Core offering

ELDA distributes validated language resources (speech corpora, written corpora, lexicons, and terminological resources) from third-party producers through the ELRA Catalogue at catalog.elra.info, selling both perpetual and evaluation licenses to academic and commercial buyers. Complementing the catalogue, ELDA delivers a service portfolio covering the full language-resource lifecycle: Identification of existing resources, Production of new corpora, Validation against standards, and Distribution with licensing/Legal Helpdesk support.

Product overview

ELDA (Evaluations and Language resources Distribution Agency) is ELRA's operational body, established in 1995 as an SME in Paris, specializing in Human Language Technologies. The core offering is the ELRA Catalogue of Language Resources - a comprehensive marketplace for speech corpora, written corpora, lexicons, and terminological resources distributed globally. Supporting this are three interconnected catalogues: the main ELRA Catalogue for commercial and research purchases, the R&D Catalogue for affordable academic use, and the Universal Catalogue for member-only resource discovery. ISLRN provides unique persistent identifiers for proper resource citation. The service portfolio spans the full language resource lifecycle: Identification (discovering resources), Production (creating new resources), Validation (ensuring quality/standards), and Distribution (licensing and delivery). Key projects include the Common European Language Data Space (LDS) platform for multilingual data sharing, the European Multilingual Web (EMW) for automated translation, and Language Technologies Solutions for speech recognition. Additional services include the License Wizard, Legal Support Helpdesk, De-identification/Anonymisation, and HLT Evaluation services. ELDA also publishes the Language Resources and Evaluation Journal and organizes the LREC conference series.

Differentiator

Problem solved

Functional benefit

Brands

  • ELRA Catalogue of Language Resources: ELRA's Language Resources Catalogue containing speech, multimodal, written lexica/corpus, and terminological resources available for purchase
  • ISLRN
  • R&D Catalogue
  • Universal Catalogue
  • LREC-COLING
  • Common European Language Data Space (LDS)

Products and services

  • ELRA Catalogue of Language Resources Comprehensive online catalogue at catalog.elra.info distributing speech corpora, multimodal corpora, written corpora, lexicons, and terminological resources on a pay-per-resource perpetual licence basis to academic and commercial buyers worldwide.
  • R&D Catalogue Academic-research subset of the ELRA Catalogue offering over 200 language resources at 500 EUR and below, targeted at universities and public research laboratories.
  • Universal Catalogue ELRA member-only service providing discovery and search-aid information about identified language resources awaiting full catalogue inclusion.
  • ISLRN (International Standard Language Resource Number) Unique persistent identifier system for language resources, jointly administered with the Lingu Data Consortium, ensuring correct identification and citation of resources in research papers, products, and evaluation benchmarks.
  • Identification Services Professional service identifying existing language resources not yet in the catalogue and negotiating with rights holders to include them in ELDA's distribution pipeline.
  • Production Services Professional service producing new language resources for natural language processing, speech systems, software localisation, language services, and electronic publishing.
  • Validation Services Standards-compliance and best-practices validation of language resources, including review by ELDA's Validation Committee, ensuring data quality and interoperability.
  • Distribution Services Licensing and distribution service covering pricing strategy, licence agreement generation, and delivery of language resources to end users globally.
  • ELRA License Wizard Online tool that helps users determine and generate appropriate licence agreements (research, commercial, evaluation) for language resources purchased via ELDA.
  • Legal Support Helpdesk Legal-assistance service for language resource procurement, including contract negotiation with providers, licensing guidance, and GDPR/data protection compliance.
  • De-identification and Anonymisation Services Personal-data protection and anonymisation service for language resources, supporting GDPR and EUDPR compliance for research and commercial users.
  • Data Management Plan (DMP) Assistance with data-management planning for language-resource projects, supporting compliance with funder and institutional DMP requirements.
  • Common European Language Data Space (LDS) Platform Three-year EU-funded platform and marketplace for collection, creation, sharing and re-use of multilingual and multimodal language data across Europe, co-developed by ELDA with DFKI, Athena RIC, and SIA Tilde.
  • Bitext Lexical Datasets and Bitext Synthetic Data 66 monolingual lexical datasets and synthetic data products added to the catalogue in July 2023, covering multiple verticals (20 verticals) in English and Spanish for intent detection and NLP applications.

Quantifiable outcome

  • 3,353 ISLRN numbers assigned to language resources worldwide
  • +2 more outcomes

Companies that use ELDA - Evaluations and Language resources Distribution Agency

Customer profile

Named customers3 records

Segments4 records

Ideal customer profiles4 records

ELDA - Evaluations and Language resources Distribution Agency technology and API

Technology

Technology focussed No

API detail

Has API
No
API docs
API detail

Core technology

AI maturity

App detail

AI capability9 records

Feature3 records

ELDA - Evaluations and Language resources Distribution Agency partnerships and signals

Strategic signal

Partnerships

15 partnerships are on record, tiered major, core and minor.

  • Lexicala by K DictionariesmajorChannel Partner/ Reseller/ Distributor · 12 October 2023Distribution agreement for Lexicala's multilingual lexical data covering 50 languages. ELDA distributes Lexicala's high-quality lexical data for language learning, machine translation, and NLP applications.
  • German Research Center for Artificial Intelligence (DFKI)coreStrategic or Co-development Partner · 19 January 2023DFKI is the coordinator of the Common European Language Data Space (LDS) project, a 3-year European Commission initiative. ELDA is one of four consortium partners working on establishing a European platform and marketplace for multilingual and multimodal language data.
  • Athena Research and Innovation Center (ILSP)coreStrategic or Co-development Partner · 19 January 2023ILSP is a consortium partner in the LDS project, responsible for developing and implementing a sustainable language data ecosystem blueprint and Language Data Space deployment.
  • SIA TildecoreStrategic or Co-development Partner · 19 January 2023SIA Tilde participates in the LDS project, focusing on implementation of multi-stakeholder data and services governance scheme and proof-of-deployment-concept projects.
  • Brno University of Technology (BUT)majorStrategic or Co-development Partner · 13 December 2022Coordinator of Language Technologies Solutions project. ELDA leads market study on Automatic Speaker Recognition solutions and coordinates collection of speech data for under-resourced European languages.
  • Tilde (Latvia)majorStrategic or Co-development Partner · 12 December 2022Coordinator of European Multilingual Web (EMW) project within Digital Europe Programme. ELDA responsible for helpdesk management supporting automated website translation solutions.
  • DatatangmajorChannel Partner/ Reseller/ Distributor · 27 October 2022Language resources distribution agreement for 67 speech resources designed to boost speech recognition. ELDA strengthens its position as worldwide distribution center while Datatang gains European market visibility.
  • Linguistic Data Consortium (LDC)coreStrategic or Co-development PartnerLong-standing partnership with LDC for distribution of language resources, joint ISLRN implementation, and collaborative projects including NETDC production/distribution and CoNLL joint distribution.
  • CLARIN ERICcoreStrategic or Co-development PartnerCollaboration agreement between ELRA and CLARIN ERIC for joint activities in language resources infrastructure and sharing.
  • INSA Rouen NormandieminorStrategic or Co-development PartnerPartnership to release the Annotated tweet corpus in Arabizi, French and English for research purposes.
  • FID LinguistikminorChannel Partner/ Reseller/ DistributorPartnership to grant licenses for written corpora to German researchers, facilitating access to language resources for the German research community.
  • IULA (Universitat Pompeu Fabra)minorStrategic or Co-development PartnerIULA has adopted ISLRN for identification of language resources.
  • Resource Management Agency (RMA)minorStrategic or Co-development PartnerRMA has adopted ISLRN for identification of language resources.
  • Vigdís International Centre of MultilingualismminorStrategic or Co-development PartnerPartnership for multilingualism and intercultural understanding initiatives.
  • ILC-ELRAminorStrategic or Co-development PartnerPartnership with ILC for language resources activities.

Scale indicators6 records

Recent moves6 records

Expansion highlights5 records

ELDA - Evaluations and Language resources Distribution Agency competitors and assessment

Company assessment

Broad incumbents

  • Defined.ai: Defined.ai (formerly DefinedCrowd) is a commercial marketplace for training data covering speech, text, and image. It competes with ELDA for the speech and NLP training data budgets of large enterprise and AI-lab buyers.
  • Appen: Appen is a publicly traded data annotation and dataset provider for AI/ML, covering text, speech, and image. While ELDA focuses on curated, validated research-grade corpora, Appen competes for the same enterprise speech/text data budgets at a much larger scale.

Direct peers

  • CLARIN ERIC: CLARIN is the European research infrastructure for language resources, with a federated catalogue of corpora and tools across European universities. ELDA has a formal collaboration agreement with CLARIN and they overlap significantly in serving academic NLP researchers with discoverable, licensed language data.
  • META-SHARE Network: META-SHARE is the federated network of European language resource repositories in which ELDA participates. It is a parallel infrastructure for sharing language data and overlaps heavily with ELDA's catalogue in target audience and resource types.
  • European Language Grid: European Language Grid is an EU-funded platform cataloguing language technology assets and resources across Europe. It is a closely adjacent infrastructure to ELDA's catalogue and serves the same multilingual European market.
  • Tilde: Tilde is a Latvian language technology company and consortium partner with ELDA on multiple EU projects (EMW, LDS). Both operate in European multilingual NLP, deliver language data services, and target enterprise/government buyers for translation and speech tools.
  • Linguistic Data Consortium (LDC): LDC at the University of Pennsylvania is the closest peer: a non-profit catalogue and distribution body for speech corpora, text corpora, and lexicons. ELDA and LDC are explicit co-administrators of ISLRN and joint distributors, and they serve the same academic and commercial NLP buyers.

Emerging players

  • Hugging Face Datasets: Hugging Face hosts a large open catalogue of NLP and speech datasets and has become the default distribution channel for ML practitioners. It overlaps with ELDA's buyer base (NLP researchers and developers) but operates on an open-source, instant-download model rather than licensed catalogue sales.
  • Datatang: Datatang is a Chinese AI data provider whose 67-resource speech catalogue is distributed by ELDA in Europe. It is comparable as a seller of speech training data, although geographic and buyer profiles differ.
  • K Dictionaries (Lexicala): Lexicala by K Dictionaries provides multilingual lexical data across 50 languages and distributes through ELDA. It is a focused peer in the lexicons/terminology sub-segment of ELDA's catalogue.

Market position

Strengths5 records

Weaknesses5 records

Competitive moat5 records

Key risks6 records

Key highlights7 records

Customer concentration

ELDA - Evaluations and Language resources Distribution Agency social profiles

Digital presence

ELDA - Evaluations and Language resources Distribution Agency compliance and trust

Trust signal

Compliance2 records

ELDA - Evaluations and Language resources Distribution Agency financial estimates

Financial estimate

Revenue estimate

Valuation estimate

ELDA - Evaluations and Language resources Distribution Agency leadership team

Management profile

Number of profiles

Profiles1 record

ELDA - Evaluations and Language resources Distribution Agency funding detail

Funding detail

Funding overview

Funding rounds

Investors

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

ELDA - Evaluations and Language resources Distribution Agency M&A and investment

M&A and investment

M&A

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about ELDA - Evaluations and Language resources Distribution Agency

What does ELDA - Evaluations and Language resources Distribution Agency do?

ELDA distributes validated language resources (speech corpora, written corpora, lexicons, and terminological resources) from third-party producers through the ELRA Catalogue at catalog.elra.info, selling both perpetual and evaluation licenses to academic and commercial buyers. Complementing the catalogue, ELDA delivers a service portfolio covering the full language-resource lifecycle: Identification of existing resources, Production of new corpora, Validation against standards, and Distribution with licensing/Legal Helpdesk support.

Is ELDA - Evaluations and Language resources Distribution Agency a public or private company?

ELDA - Evaluations and Language resources Distribution Agency is a private company. It is classified as management employee owned and is currently operating.

When was ELDA - Evaluations and Language resources Distribution Agency founded?

ELDA - Evaluations and Language resources Distribution Agency was founded in 1995. It employs 51 to 100 people.

Where is ELDA - Evaluations and Language resources Distribution Agency based?

ELDA - Evaluations and Language resources Distribution Agency is headquartered in Paris, France, in the Europe region.

How does ELDA - Evaluations and Language resources Distribution Agency make money?

Three revenue lines are on record. Language Resources Sales are the primary driver. The others are membership Fees and licensing Services.

Who are ELDA - Evaluations and Language resources Distribution Agency's main competitors?

Broad incumbents on record are Defined.ai and Appen. Direct peers are CLARIN ERIC, META-SHARE Network, European Language Grid, Tilde and Linguistic Data Consortium (LDC). Emerging players are Hugging Face Datasets, Datatang and K Dictionaries (Lexicala).

Does ELDA - Evaluations and Language resources Distribution Agency have an API?

No public API is recorded for ELDA - Evaluations and Language resources Distribution Agency.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales