Pangeanic
Pangeanic provides multilingual AI data, model alignment, and sovereign deployment infrastructure — including datasets, RLHF, machine translation, anonymization, and small language model customization — for AI labs, enterprises, and regulated public sector organizations across 84 languages.
- Company typePrivate
- Founded2000
- HeadquartersValencia, Spain
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Pangeanic does
Pangeanic is a privately held multilingual AI infrastructure company headquartered in Valencia, Spain, founded in 2000 as a language service provider and repositioned since 2017 around AI data and language technologies. The company builds and operates a stack spanning the full multilingual AI lifecycle: the ECO Intelligence Platform for orchestrating secure translation, retrieval, anonymization, and enterprise APIs; the PECAT data annotation platform for managing multilingual annotation, human review, RLHF, and evaluation pipelines; proprietary Deep Adaptive AI Translation (DAAIT) and Machine Translation Quality Estimation (MTQE) engines; small/task-specific language model customization; and sovereign deployment infrastructure supporting private cloud, on-premises, and air-gapped environments. Underlying these products is a sizable data asset: 10 billion+ translation alignments across 84 languages, a 24+ EU language production base, and 506 neural MT engines built through the EU-funded NTEU program, alongside ISO 27001, 9001, 13485, 17100 and 18587 certifications and GDPR-aligned multilingual anonymization.
Pangeanic serves three primary customer segments: AI labs and model builders requiring legally usable, domain-relevant multilingual training data; enterprise AI teams operationalizing AI workflows with translation, anonymization, quality estimation, and integration needs; and regulated/public sector organizations requiring sovereign, auditable, air-gapped AI deployments. Named customers include the Barcelona Supercomputing Center (data and alignment for the Salamandra sovereign LLM), Spain's Tax Agency (AEAT), Veritone and DoD Iron Bank, EFE News Agency, Amazon (multilingual idiomatic corpus), the Government of Catalonia, and Europeana. Revenue is generated through a mix of dataset licensing, bespoke data collection, AI data operations and RLHF professional services, subscription-based machine translation, and managed sovereign deployment — all delivered via quote-based enterprise contracts, with no channel partner model evident.
The business is founder-led (CEO Manuel Herranz) with no disclosed institutional venture or private equity backing, supplemented by EU/CEF research grants and Spanish public innovation funding (most notably a €435,083 CDTI Innoglobal 2025 award scoring 85/100). The company operates from Valencia with additional offices in London, Tokyo, Boston, New York, Hong Kong, and Shanghai, and is recognized by Gartner as a Representative Vendor in Conversational AI and Data Masking/Synthetic Data.
Pangeanic firmographics
Firmographics- Name
- Pangeanic
- Legal name
- Pangeanic B.I. Europa S.L.
- Website
- https://pangeanic.com
- Company type
- Private
- Founded year
- 2000
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- Pangeanic provides multilingual AI data, model alignment, and sovereign deployment infrastructure — including datasets, RLHF, machine translation, anonymization, and small language model customization — for AI labs, enterprises, and regulated public sector organizations across 84 languages.
- Ownership category
- akta.pro rank
Pangeanic industry classification
Industry- Product category
- Multilingual AI Infrastructure
- NAICS
- Computer Systems Design and Related Services (5415), Software Publishers (5132), Computer Systems Design and Related Services (54151)
- SIC
- Services-Computer Programming, Data Processing, Etc. (7370), Services-Computer Integrated Systems Design (7373), Services-Prepackaged Software (7372)
- akta.pro primary industry
- End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management) (HDAEANAA)
- akta.pro secondary industries
- Enterprise Foundation Model Integration & APIs (Connectors, Governance, Deployment) (HDAAACAO), Generative AI & LLM Solutions Services (RAG, Agents, Copilots) (BPAEAHAG), Enterprise AI Data & Knowledge Platforms (Vector Databases, Knowledge Graphs) (HDAEANAH), Responsible AI, Security & Privacy Platforms (Safety, Guardrails, PII) (HDAEANAG)
Keywords
Where Pangeanic is headquartered
LocationHeadquarters
- HQ city
- Valencia
- HQ country
- Spain
- HQ region
- Europe
Offices8 records
Markets served
Pangeanic business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Operations, Infrastructure, Marketing or Sales
Revenue model
- AI Dataset Licensing: License of existing multilingual text, parallel corpora, speech and audio, image, video, OCR, multimodal data and evaluation sets. Ready-to-license datasets for faster procurement.
- Bespoke Data Collection: Custom data programs designed around specific languages, domains, demographic requirements, format, consent model and quality thresholds when existing datasets don't match requirements.
- AI Data Operations Services: Sourcing, collection, annotation, metadata, human review, RLHF and model evaluation operations. Includes annotation design, expert-review workflows, benchmark building and continuous evaluation.
- Machine Translation Services: Enterprise machine translation including adaptive translation, terminology control, quality estimation, secure APIs, document processing for PDF, Word, PowerPoint, Excel and scanned documents.
- Sovereign AI Deployment: On-premises and air-gapped deployment services, private cloud infrastructure, and controlled deployment options for regulated environments requiring data residency.
- Model Alignment & RLHF: Human feedback, preference ranking, policy-aware review, and alignment workflows that shape model behavior. Includes evaluation datasets and benchmarking services.
Go-to-market motion3 records
Distribution channels4 records
Marketing channels8 records
Pangeanic product offering
Product offeringCore offering
Pangeanic provides multilingual AI infrastructure combining licensed datasets, human-feedback and RLHF model alignment services, secure machine translation (DAAIT, MTQE), anonymization, and sovereign deployment on private cloud, on-premises, or air-gapped environments. Its products — including the ECO Intelligence Platform, PECAT annotation platform, ECOChat, and small language model customization — enable enterprises, AI labs, and public administrations to build and operate multilingual AI systems under their own governance rather than relying on third-party APIs.
Product overview
Pangeanic provides a unified multilingual AI infrastructure combining data, model alignment, and sovereign deployment capabilities. The core platform is the ECO Intelligence Platform, which orchestrates secure translation, quality estimation, anonymization, multilingual RAG, and enterprise APIs. Supporting products include PECAT for data annotation and RLHF workflows, ECOChat for virtual assistant functionality, and specialized machine translation offerings (Machine Translation, Deep Adaptive AI Translation, MTQE, Enterprise AI Document Translator). Small Language Model Customization enables task-specific deployments, while Data Masking and Anonymization provide privacy controls. The product suite spans Datasets for AI, Off-the-Shelf Training Data, Model Alignment and RLHF, Evaluation and AI QA, AI Data Operations, and Sovereign AI Systems. Together these components address the full pipeline from multilingual data sourcing through production deployment under organizational governance.
Differentiator
Problem solved
Functional benefit
Brands
- PangeaMT: Machine translation platform developed by Pangeanic for custom-built enterprise machine translation solutions.
- PECAT
- ECO Intelligence Platform
- ECO LLM
- ECOChat
- Deep Adaptive AI Translation (DAAIT)
- MTQE
Products and services
- ECO Intelligence Platform Central orchestration layer for multilingual AI operations connecting secure translation, quality estimation, anonymization, multilingual RAG, and enterprise APIs, supporting private cloud, on-premises, and air-gapped deployments for controlled multilingual knowledge discovery and document processing. Targeted at AI labs, enterprise AI teams, and regulated/public-sector organizations.
- PECAT Data Annotation Platform Operational data annotation and orchestration platform for managing multilingual and multimodal AI data workflows, providing human review, evaluation pipelines, RLHF workflows, traceability, and quality control across the AI data lifecycle from collection to delivery. Targeted at AI labs and enterprise AI teams preparing data for foundation models.
- ECOChat Multilingual virtual AI assistant powered by customer data, enabling secure conversational AI with grounded responses across languages using controlled knowledge sources and enterprise data. Targeted at enterprises and public administrations needing controlled multilingual chat.
- Machine Translation Neural machine translation engines covering EU and global languages, supporting on-premises, private cloud, and air-gapped deployment for secure enterprise and public-sector document translation workflows. Targeted at organizations requiring controlled translation infrastructure.
- Deep Adaptive AI Translation (DAAIT) Adaptive machine translation using client terminology, translation memories, and domain resources for tone, terminology, and domain control, incorporating multi-agent translation, post-editing automation, and specialized translation agent workflows. Targeted at enterprises with domain-specific translation needs.
- Machine Translation Quality Estimation (MTQE) Quality estimation system providing confidence scores for MT output, automatic corrective loops, and review routing by quality threshold, enabling efficient human review allocation and quality control in high-volume translation operations. Targeted at translation operations teams.
- Enterprise AI Document Translator AI-powered document translation for PDF, Word, PowerPoint, Excel, and scanned documents, supporting enterprise-scale multilingual document workflows with terminology control and quality assurance. Targeted at enterprises with large-scale multilingual documentation.
- Small Language Model Customization Task-specific language model selection, fine-tuning, and deployment for defined enterprise and public-sector workflows, with model adaptation to organizational terminology, proprietary knowledge, policies, and infrastructure requirements and lower inference costs. Targeted at organizations needing sovereign, cost-efficient custom models.
- Data Masking and Anonymization Multilingual data anonymization and PII masking for GDPR compliance in public-sector, medical, and legal domains, including named entity recognition, de-identification, pseudo-anonymization, and speech/video data masking for privacy protection. Targeted at regulated and public-sector organizations.
- Datasets for AI Licensed and bespoke multilingual datasets including parallel corpora, speech/audio, image/video, OCR, multimodal data, and evaluation sets across 24+ EU languages and global coverage, supporting OSINT, entity analysis, classification, retrieval, and model evaluation. Targeted at AI labs and model builders.
- Off-the-Shelf Training Data Ready-to-license commercial datasets for faster AI procurement, including multilingual text, parallel corpora, speech, audio, image, video, and multimodal data with quality controls and provenance documentation. Targeted at AI labs and enterprises needing immediate dataset access.
- Model Alignment and RLHF Human feedback, preference ranking, policy-aware review, and alignment workflows that shape model behavior, including expert validation, disagreement resolution, preference optimization, and continuous evaluation for model release cycles. Targeted at AI labs and enterprises aligning models for responsible behavior.
- Evaluation and AI QA Benchmark design, multilingual QA, regression testing, scoring, and validation frameworks for dependable AI release cycles, providing gold-standard evaluation sets, model diagnostics, and continuous alignment measurement. Targeted at AI labs and enterprises running production AI.
- AI Data Operations Full-cycle AI data services including bespoke data collection, annotation, metadata structuring, human review, RLHF, and evaluation operations, managing multilingual annotation, quality control, and delivery workflows for production-ready AI assets. Targeted at AI labs and enterprises operationalizing AI data pipelines.
- Sovereign AI Systems On-premises, private cloud, and air-gapped AI deployment infrastructure for organizations requiring control over data, models, and operational policy, supporting secure multilingual systems with auditable workflows and data residency compliance. Targeted at regulated, defense, and public-sector organizations.
Quantifiable outcome
- 10Bn+ translation alignments across 84 languages
- +2 more outcomes
Companies that use Pangeanic
Customer profileNamed customers8 records
Segments3 records
Ideal customer profiles3 records
Pangeanic technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- Yes
- API docs
- API detail
Core technology
AI maturity
App detail
Integration1 record
AI capability14 records
Feature8 records
Pangeanic partnerships and signals
Strategic signalPartnerships
13 partnerships are on record, tiered strategic, minor, core and notable.
- INTRAS (University of Valencia)strategicLeads DS4M Mediterráneo (Data Space for Mobility) project. Pangeanic contributes data engineering, quality standardization, and AI-driven analytics for secure, sovereign mobility data sharing.
- IRTICminorResearch technology institute collaborating in DS4M Mediterráneo project for mobility data space development.
- ITS SpainminorIntelligent Transport Systems association supporting DS4M Mediterráneo ecosystem for mobility data space implementation.
- Barcelona Supercomputing Center (BSC)coreCollaboration on multilingual AI projects including data annotation, RLHF workflows, and evaluation datasets for large language model training. Contributed to BSC's Language Technologies Unit work on language models, translation and NLP research, including the Salamandra models.
- KantanMTcoreConsortium partner in NTEU project to build the largest neural machine translation engine farm (506 engines) for all EU language combinations.
- TildecoreConsortium partner in NTEU project for European machine translation infrastructure development.
- AmazonnotableAmazon multilingual corpus project building a multilingual corpus of idiomatic expressions across languages and cultural contexts. Executed through coordinated workflows between internal teams and external linguists.
- European Commission / CEFcoreMultiple CEF-funded projects including NTEU, NEC TM, MAPA, iADAATPA/MT-Hub, Europeana Translate, Culture Chatbot, J-Ark, Jewish History Tours, Europeana XX Century of Change.
- Digital Europe ProgrammecorePartner in MOSAIC Media (multilingual media platform) and AI4Culture (AI capacity-building for cultural heritage institutions) under DIGITAL Europe Programme.
- European Language Equality (ELE/ELE2)coreContribution to European language equality work including large speech corpus generation for the languages of Spain using data augmentation.
- ValgrAInotableBoard of trustees membership in Valencian Graduate School and Research Network of Artificial Intelligence for advancing AI innovation.
- Everis/NTTDatanotableConsortium partner in iADAATPA/MT-Hub project for secure automatic translation platform for EU public administrations.
- EU Next Generation / DS4M MediterráneostrategicParticipation in mobility data space initiative funded by EU Next Generation and Spanish Recovery, Transformation and Resilience Plan. Project TSI-100120-2024-9.
Scale indicators7 records
Pangeanic competitors and assessment
Company assessmentDirect peers
- Lionbridge AI (DataForce): Lionbridge's DataForce division offers multilingual data collection, annotation and AI evaluation services. Closely comparable to Pangeanic's dataset, annotation and RLHF services across 80+ languages.
- Tilde: European language technology company providing MT, speech and AI localization services, and a long-time NTEU consortium partner with Pangeanic. Competes directly in EU-language MT, public-sector deployment and sovereign AI tooling.
- Lilt: Enterprise adaptive machine translation platform with integrated human-in-the-loop workflows. Closely aligned with Pangeanic's DAAIT and PECAT offering for organizations that need controlled, terminology-aware translation at enterprise scale.
- Appen: Global leader in training data for AI, offering data collection, annotation, RLHF and evaluation across languages and modalities. Directly comparable to Pangeanic's AI Data Operations, PECAT annotation workflows and multilingual dataset offerings, though at significantly larger scale.
- Welocalize: Global language services provider with a dedicated AI services practice (Welocalize AI) covering data annotation, MT evaluation and model evaluation. Direct competitor in the multilingual AI data and evaluation space Pangeanic targets with PECAT and AI Data Operations.
- RWS Group: Language services and IP services group with a large Language Technology arm (MT, localization, Trados, Language Weaver). Overlaps with Pangeanic's adaptive MT, MTQE and enterprise translation offering, and addresses similar regulated-industry buyer profiles.
- Unbabel: AI-powered translation platform combining MT with human post-editing and a quality estimation layer (Widn.AI). Directly comparable to Pangeanic's DAAIT adaptive translation and MTQE positioning, and shares a similar enterprise + public-sector GTM.
- TransPerfect: One of the largest language services and AI data companies globally, with translation, localization, AI data services and secure enterprise deployments. Comparable across machine translation, professional services for regulated buyers and on-prem enterprise language technology.
- Scale AI: AI data platform specializing in high-quality annotation, RLHF and evaluation for foundation model builders. Overlaps with Pangeanic's Model Alignment, RLHF and Evaluation & AI QA offerings, although Scale's primary focus has been English-centric.
Broad incumbents
- AWS (Amazon Translate / Bedrock): Hyperscaler offering managed MT, foundation models and data services as part of a broad cloud AI portfolio. Competes with Pangeanic's MT, small-language-model customization and sovereign deployment offerings, particularly where buyers accept cloud-resident AI.
Market position
Strengths5 records
Weaknesses5 records
Key risks6 records
Key highlights7 records
Customer concentration
Pangeanic social profiles
Digital presencePangeanic compliance and trust
Trust signalCompliance4 records
Pangeanic financial estimates
Financial estimateRevenue estimate
Valuation estimate
Pangeanic leadership team
Management profileNumber of profiles
Profiles7 records
Pangeanic funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Pangeanic M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Pangeanic
What does Pangeanic do?
Pangeanic provides multilingual AI infrastructure combining licensed datasets, human-feedback and RLHF model alignment services, secure machine translation (DAAIT, MTQE), anonymization, and sovereign deployment on private cloud, on-premises, or air-gapped environments. Its products — including the ECO Intelligence Platform, PECAT annotation platform, ECOChat, and small language model customization — enable enterprises, AI labs, and public administrations to build and operate multilingual AI systems under their own governance rather than relying on third-party APIs.
Is Pangeanic a public or private company?
Pangeanic is a private company. It is classified as founder individual operated bootstrapped and is currently operating.
When was Pangeanic founded?
Pangeanic was founded in 2000. It employs 11 to 50 people.
Where is Pangeanic based?
Pangeanic is headquartered in Valencia, Spain, in the Europe region.
How does Pangeanic make money?
Six revenue lines are on record. AI Dataset Licensing is the primary driver. The others are bespoke Data Collection, AI Data Operations Services, machine Translation Services, sovereign AI Deployment and model Alignment & RLHF.
Who are Pangeanic's main competitors?
Direct peers on record are Lionbridge AI (DataForce), Tilde, Lilt, Appen, Welocalize, RWS Group, Unbabel, TransPerfect and Scale AI. AWS (Amazon Translate / Bedrock) is listed as a broad incumbent.
Does Pangeanic have an API?
Yes. Pangeanic offers API-based services primarily for enterprise and public sector workflows. The NEC TM (National European Central Translation Memory) platform provides API-based retrieval with CAT tool independence. The iADAATPA/MT-Hub system includes routing, language detection, and scalable public sector translation infrastructure with connectors. The ECO Intelligence Platform offers secure APIs and enterprise integration for connecting repositories, portals, document systems and internal applications. APIs support multilingual machine translation, MTQE, anonymization, quality routing, and document processing workflows.
What industry is Pangeanic in?
Pangeanic's product category is Multilingual AI Infrastructure. Its primary akta.pro industry code is HDAEANAA, End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management), with a secondary code of HDAAACAO, Enterprise Foundation Model Integration & APIs (Connectors, Governance, Deployment). Its NAICS code is 5415 and its SIC code is 7370.