pleias
Pleias is a Paris-based AI company that builds open, rights-cleared data infrastructure, synthetic data tooling, and specialized language models for enterprises and researchers needing on-premise, domain-specific AI deployments.
- Company typePrivate
- Founded2023
- HeadquartersParis, France
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What pleias does
Pleias is a French AI company building the data infrastructure layer for specialized and sovereign AI. Founded in 2023 and headquartered at Station F in Paris, the company develops three core products: Common Corpus, the largest open, rights-cleared and provenance-based dataset for LLM pre-training; Synth, a synthetic data generation platform that simulates expert-level reasoning to produce domain-specific training data and runs on-premise; and Stratum, AI-native tooling that turns siloed documents into structured, compliant data assets for agentic workflows with a built-in PII firewall. The portfolio also includes specialized datasets and models such as Telco Common Corpus (10B+ tokens, co-developed with GSMA), CommonLingua (334-language identification), and the Sillon 600M-parameter model deployed at RATP.
Underlying technology centers on a synthetic-data generation pipeline that uses backtranslation with fine-tuned models and synthetic reasoning traces to produce training data for vertical applications. Pleias's core thesis is that smaller, specialized models trained on rights-cleared, domain-specific data can match or exceed proprietary LLMs at a fraction of the parameter count, and they can be deployed fully on-premise to meet sovereignty and privacy requirements. The technology has been validated through peer-reviewed research (ICLR 2025 Oral for Common Corpus, ACL Findings for synthetic-data sentiment analysis) and through production deployments such as Sillon at RATP's sovereign infrastructure on Scaleway.
Pleias operates a hybrid commercial model: enterprise teams access the full data infrastructure stack via quote-based annual subscriptions booked through direct demos, while research datasets and lighter models are distributed as freemium through HuggingFace and GitHub. Customers and partners include RATP (transportation), SpineDAO (healthcare), GSMA (telecommunications), NVIDIA (AI Factory France and Nemotron Personas co-development), Wikimedia Foundation, Mozilla, AI Alliance, Scaleway, AT&T, AMD, and Huawei Technologies France. Co-founder Anastasia Stasenko has positioned Pleias publicly around the idea that domain expertise and curated data are sustainable competitive advantages in the AI industry.
pleias firmographics
Firmographics- Name
- pleias
- Legal name
- Pleias
- Website
- https://pleias.fr
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- Pleias is a Paris-based AI company that builds open, rights-cleared data infrastructure, synthetic data tooling, and specialized language models for enterprises and researchers needing on-premise, domain-specific AI deployments.
- Ownership category
- akta.pro rank
pleias industry classification
Industry- Product category
- AI Data Infrastructure
- NAICS
- Software Publishers (51321), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (5182), Computing Infrastructure Providers, Data Processing, Web Hosting, and Related Services (518)
- SIC
- Services-Prepackaged Software (7372), Services-Computer Processing & Data Preparation (7374)
- akta.pro primary industry
- Synthetic Data & Data Augmentation for Foundation Models (HDAAACAK)
- akta.pro secondary industries
- Dataset Discovery, Marketplaces & Licensing (HDAAALAE), Data Pipelines for GenAI (Curation, Filtering, Deduplication, Copyright) (HDAAACAJ), Enterprise AI Data & Knowledge Platforms (Vector Databases, Knowledge Graphs) (HDAEANAH)
Keywords
Where pleias is headquartered
LocationHeadquarters
- HQ city
- Paris
- HQ country
- France
- HQ region
- Europe
Offices1 record
Markets served
pleias business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Infrastructure, Marketing or Sales, Operations
Revenue model
- Enterprise Data Infrastructure: Subscription-based access to Pleias' data infrastructure products (Stratum, Synth, Common Corpus) for enterprise teams. Provides the data layer behind models that beat larger ones at lower cost.
- Open Dataset Access: Open datasets available freely for researchers and developers through HuggingFace platform, with potential for premium support or customization services.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Subscription | Annual | Enterprise Teams - Full data infrastructure access |
Go-to-market motion2 records
Distribution channels4 records
Marketing channels7 records
pleias product offering
Product offeringCore offering
Pleias builds the data layer for AI by providing three core products: Synth, which generates synthetic training data through simulated expert-level reasoning; Common Corpus, the largest open, rights-cleared dataset for LLM pre-training across government, legal, scientific, and multilingual sources; and Stratum, which transforms siloed enterprise documents into structured, compliant data assets for agentic AI workflows. These products are sold via enterprise subscriptions and made available on-premise, while open datasets are distributed freely through HuggingFace.
Product overview
Pleias is a data infrastructure company that builds the data layer for AI, offering three core products: Synth (synthetic data generation for AI agents), Common Corpus (rights-cleared open training datasets for LLMs), and Stratum (AI-native data structuring tooling for agentic workflows). These products are designed to work together or independently, with Synth generating specialized training data, Common Corpus providing foundational open data, and Stratum enabling enterprises to structure their own documents for RAG pipelines and MCP servers. The company also offers specialized domain-specific products like Telco Common Corpus for telecommunications and has developed language identification and synthetic persona models.
Differentiator
Problem solved
Functional benefit
Brands
- Common Corpus: Fully Open Data for AI - a rights-cleared and provenance-based dataset for LLMs including government records, legal archives, scientific literature, and multilingual sources.
- Synth
- Stratum
- Telco Common Corpus
- CommonLingua
Products and services
- Synth Synthetic data generation platform that simulates expert-level reasoning and domain-specific processes to generate training data for AI agent specialization. Handles cold-start problems, covers long-tail edge cases through engineered simulations, and runs entirely on-premise to keep sensitive data under customer control. Designed for enterprise teams training specialized models on proprietary or compliance-sensitive workflows.
- Common Corpus Fully open, rights-cleared and provenance-based dataset for LLM pre-training, covering government records, legal archives, scientific literature, and multilingual sources. Available on HuggingFace for direct integration into models, RAG pipelines, and MCP servers. The largest open corpus for LLM pre-training with verified legal compliance for data security regulations.
- Stratum AI-native tooling that turns siloed enterprise documents into a single structured, compliant data asset for agentic AI workflows. Includes a built-in privacy firewall that handles PII before anything leaves the secure zone, and is deployable fully on-premise for enterprises with strict data sovereignty requirements.
- Telco Common Corpus 10-billion+ token open telecommunications knowledge library assembled from high-quality open data sources including 3GPP documents and US/EU patents. Designed for telecom operators, vendors, and stakeholders seeking to build AI models focused on the telecom sector.
- CommonLingua Open-source language identification model covering 334 languages including 61 African languages across eight African language families. Operates on UTF-8 byte sequences and achieves 83% accuracy, addressing digital inclusion gaps in African markets and underrepresented-language AI applications.
- Sillon Specialized 600-million-parameter model for RATP (Parisian Subway Operator) that detects and interprets safety signals in Parisian Subway users' messages. Combines a fully synthetic training pipeline and is designed for on-premise deployment, beating closed models 200x its size in production at RATP's sovereign infrastructure on Scaleway.
- Nemotron-Personas-Belgium Statistically grounded synthetic persona dataset covering the Belgian population at the level of regions, language communities, and communes. Released in collaboration with NVIDIA as the second European dataset in the Nemotron Personas series.
- Nemotron-Personas-France Synthetic persona dataset covering the French population, the first dataset in the Nemotron Personas series developed in collaboration with NVIDIA.
Quantifiable outcome
- 600M parameter specialized model beats closed models 200x its size in production
- +3 more outcomes
Companies that use pleias
Customer profileNamed customers4 records
Segments5 records
Ideal customer profiles4 records
pleias technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability4 records
Feature7 records
pleias partnerships and signals
Strategic signalPartnerships
Eleven partnerships are on record, tiered flagship, core and major.
- GSMAflagshipGSMA partnered with French AI startup Pleias to launch the Telco Common Corpus (TCC), a 10-billion-token open telecommunications knowledge library for training AI models. This is part of the GSMA Open Telco AI initiative launched at Mobile World Congress 2026 to develop telco-grade AI models through industry collaboration. GSMA and Pleias also released CommonLingua, an open-source language identification model covering 334 languages including 61 African languages.
- NVIDIAflagshipPleias is part of AI Factory France (AI2F), led by GENCI, in partnership with NVIDIA to provide startups with streamlined access to NVIDIA AI supercomputers. At VivaTech 2026, Pleias and NVIDIA released Nemotron-Personas-Belgium, a statistically grounded synthetic persona dataset covering the Belgian population. This follows the Nemotron-Personas-France announcement in March 2026.
- AT&TcoreAT&T is a founding supporter of the GSMA Open Telco AI initiative alongside Pleias, contributing open telco models to address gaps where general-purpose AI models struggle with network data interpretation and standards documentation.
- AMDcoreAMD is a founding supporter of the GSMA Open Telco AI initiative alongside Pleias, contributing compute capacity to the initiative focused on developing AI models tailored for telecom requirements.
- Huawei Technologies FrancecoreHuawei Technologies France is listed as a supporter of the GSMA Open Telco AI initiative, which includes Pleias as a key data partner for developing telco-grade AI models.
- Wikimedia FoundationcoreWikimedia Foundation is listed as a partner in Pleias' Partners & Ecosystem section. The partnership involves licensing Wikipedia content through Wikimedia Enterprise platform for AI applications.
- MozillacoreMozilla is listed as a partner in Pleias' Partners & Ecosystem section, indicating alignment on open-source AI values and ecosystem collaboration.
- AI AlliancecoreAI Alliance is listed as a partner in Pleias' Partners & Ecosystem section, representing collaboration with a major industry consortium focused on open-source AI development.
- ScalewaycoreScaleway is listed as a partner and infrastructure provider. Pleias' Sillon model for RATP is deployed on Scaleway's sovereign infrastructure, demonstrating integration with Scaleway's cloud and AI supercomputer services.
- RATPflagshipPleias trained a 600-million-parameter specialized model (Sillon) for RATP to detect and interpret safety signals in Parisian Subway users' messages. The model is now in production at RATP's sovereign infrastructure on Scaleway after only three months of development.
- SpineDAOmajorPleias and SpineDAO are partnering to build AI systems that safely scale expert spine care for back pain, the world's leading cause of disability. The project tests whether small, structured-reasoning models can outperform large generic LLMs in high-stakes clinical settings.
Scale indicators7 records
Recent moves6 records
Expansion highlights6 records
pleias competitors and assessment
Company assessmentDirect peers
- Mostly AI: Mostly AI offers a synthetic data platform focused on enterprise privacy-preserving data generation, similar to Pleias' Synth. Both target regulated industries needing on-premise deployment and structured synthetic data, making it a direct European peer in synthetic data tooling.
- Surge AI: Surge AI is a data labeling and evaluation platform for frontier AI models, with growing synthetic data capabilities. Comparable to Pleias in high-quality training data generation, but operating at a different scale and with foundation model lab customers rather than European sovereign buyers.
- Scale AI: Scale AI is a leading data labeling and synthetic data platform serving foundation model labs and enterprises. Directly comparable to Pleias in synthetic data generation and data-centric AI services, but at vastly larger scale and with a more commercial, US-centric customer base.
- Snorkel AI: Snorkel AI provides data-centric AI tooling including programmatic labeling and synthetic data for enterprise model training. Comparable to Pleias in enabling specialized enterprise models with curated data, though focused on labeling workflows and US enterprise sales.
- Gretel AI (Snowflake): Gretel provides a synthetic data platform for AI training and testing, with strong enterprise privacy and compliance positioning. Acquired by Snowflake, it directly competes with Pleias' Synth product in synthetic data for regulated industries, though with a broader US enterprise focus.
Emerging players
- Datasaur: Datasaur offers an NLP data labeling and LLM evaluation platform serving enterprise teams. Comparable to Pleias in the labeled training data layer for language models, but more focused on labeling workflows than synthetic data generation or rights-cleared open corpora.
- Encord: Encord offers data labeling, curation, and evaluation tooling for multimodal AI training. Comparable to Pleias in the data preparation layer for AI, though with broader modality support (image, video) and less emphasis on rights-cleared text datasets or on-premise sovereign deployment.
- Kili Technology: Kili Technology provides a data labeling and curation platform for AI training, headquartered in Paris like Pleias. A regional peer in the data preparation layer, with overlap in enterprise data tooling for AI but less emphasis on synthetic data generation and rights-cleared open datasets.
- Lightly: Lightly provides data curation and selection tooling for AI training, helping teams identify the most valuable samples in large datasets. Adjacent to Pleias' Common Corpus curation pipeline and Stratum structuring offering, focused more on sample selection than rights-cleared provenance.
Broad incumbents
- Hugging Face: Hugging Face is the dominant open AI platform hosting models, datasets, and spaces, and is the primary distribution channel for Pleias' Common Corpus and CommonLingua releases. A broad incumbent rather than direct competitor, but its dataset hosting and enterprise offerings shape the market Pleias operates in.
Market position
Strengths5 records
Weaknesses5 records
Competitive moat5 records
Key risks5 records
Key highlights6 records
Customer concentration
pleias social profiles
Digital presencepleias financial estimates
Financial estimateRevenue estimate
Valuation estimate
pleias leadership team
Management profileNumber of profiles
Profiles1 record
pleias funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
pleias M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about pleias
What does pleias do?
Pleias builds the data layer for AI by providing three core products: Synth, which generates synthetic training data through simulated expert-level reasoning; Common Corpus, the largest open, rights-cleared dataset for LLM pre-training across government, legal, scientific, and multilingual sources; and Stratum, which transforms siloed enterprise documents into structured, compliant data assets for agentic AI workflows. These products are sold via enterprise subscriptions and made available on-premise, while open datasets are distributed freely through HuggingFace.
Is pleias a public or private company?
pleias is a private company. It is classified as unknown and is currently operating.
When was pleias founded?
pleias was founded in 2023. It employs 1 to 10 people.
Where is pleias based?
pleias is headquartered in Paris, France, in the Europe region.
How does pleias make money?
Two revenue lines are on record. Enterprise Data Infrastructure is the primary driver. The others are open Dataset Access.
Who are pleias's main competitors?
Direct peers on record are Mostly AI, Surge AI, Scale AI, Snorkel AI and Gretel AI (Snowflake). Emerging players are Datasaur, Encord, Kili Technology and Lightly. Hugging Face is listed as a broad incumbent.
Does pleias have an API?
No public API is recorded for pleias.
What industry is pleias in?
pleias's product category is AI Data Infrastructure. Its primary akta.pro industry code is HDAAACAK, Synthetic Data & Data Augmentation for Foundation Models, with a secondary code of HDAAALAE, Dataset Discovery, Marketplaces & Licensing. Its NAICS code is 51321 and its SIC code is 7372.