Pareto AI
Pareto AI is a US-based expert data infrastructure company that supplies frontier AI labs, government safety institutes, and academic researchers with vetted specialist human training data and verifier engineering for AI post-training, delivered through its Forte expert workforce platform.
- Company typePrivate
- Founded2020
- HeadquartersWilmington, United States
- Headcount101–250
- GTM typeB2B
- OfferingSoftware
What Pareto AI does
Pareto AI (Pareto, Inc.) is a privately held Delaware corporation founded in 2020, registered at 4023 Kennett Pike #52468, Wilmington, DE, with operations spanning the United States, United Kingdom, Philippines, the European Economic Area, and Switzerland. The company supplies frontier AI laboratories (Anthropic, xAI, DeepMind, ByteDance, Amazon), government AI safety institutes (UK AISI), and academic research consortia (MATS, UCL) with expert human training data and verifier engineering services for RL post-training. Revenue is generated through custom-quote, project-based engagements with monthly subscription billing terms, sourced via a high-touch enterprise sales motion routed through a Typeform 'Request a Project' form and a dedicated [email protected] inbox. The business was founded by Phoebe Yao, a Thiel Fellow and Forbes 30 Under 30 honoree, and operates remote-first.
The core product is the Forte platform (forte.pareto.ai), a proprietary expert workforce system that recruits, qualifies, and matches vetted domain specialists — physicians, lawyers, PhDs, engineers, finance professionals, ML developers, and linguists — to client AI training projects. Underlying technology combines Forte's task-specification and direct researcher-to-worker communication tooling with a proprietary verifier engineering methodology (a five-property quality framework — consistency, calibration, coverage, robustness, auditability — paired with a taxonomy of ten common verifier failure modes) that translates expert judgment into RL post-training reward signals via programmatic checks, LLM-as-judge, and agentic graders. The company publishes open research benchmarks (LEAPBench for in-context learning efficiency across 55 sequential optimization tasks; AttuneBench for emotional intelligence grounded in the Mayer-Salovey-Caruso model with 200 conversations and 50,000+ annotations) and has fine-tuned open-weight models, notably lifting Llama-3.1-8B from 77% to 96% accuracy on a harmful-advice classifier co-developed with the UK AISI and deployed live to protect 2,302 study participants. Cloud infrastructure runs on AWS and Vercel with Stripe for billing, Auth0 for identity, and Intercom and Sendgrid for communications.
The go-to-market combines direct enterprise sales, a community-led supply-side funnel through Forte that recruits experts across medicine, law, STEM, finance, software/ML, and humanities, and co-development partnerships with research consortia and government safety institutes that double as flagship case studies. Pareto explicitly targets the approximately 90% of expert work in healthcare, legal, accounting, software engineering, and data science that RLVR cannot deterministically verify, positioning verifier engineering methodology as its durable moat against commodity data-labeling competitors.
Pareto AI firmographics
Firmographics- Name
- Pareto AI
- Legal name
- Pareto Inc.
- Website
- https://pareto.ai
- Company type
- Private
- Founded year
- 2020
- Operating status
- Operating
- Headcount range
- 101–250 employees
- Short description
- Pareto AI is a US-based expert data infrastructure company that supplies frontier AI labs, government safety institutes, and academic researchers with vetted specialist human training data and verifier engineering for AI post-training, delivered through its Forte expert workforce platform.
- Ownership category
- akta.pro rank
Pareto AI industry classification
Industry- Product category
- Expert AI Training Data Platform
- NAICS
- Scientific Research and Development Services (5417), Professional, Scientific, and Technical Services (541)
- SIC
- Services-Computer Programming Services (7371)
- akta.pro primary industry
- Bias, Fairness & Representativeness Testing for Datasets (HDAAALAK)
- akta.pro secondary industry
- Safety & Alignment Evaluation (red-teaming, harmful capability testing) (HDAAAMAL)
Keywords
Where Pareto AI is headquartered
LocationHeadquarters
- HQ city
- Wilmington
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Pareto AI business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Operations, Marketing or Sales
Revenue model
- Custom AI Training Data Projects: Clients (frontier AI labs, research institutions, government safety institutes) propose AI training projects; Pareto creates a budget, recruits and supervises expert contractors, and delivers verified training data. Revenue is project-based, scoped per initiative (e.g., debate judgments, harmful-advice classification, expert-labeled RL tasks), with bespoke pricing rather than off-the-shelf tiers.
- Subscription / Recurring Service Plans: Terms of Use reference subscription plans with monthly billing cadence, cancellation rights at end of current paid term, prorated refunds upon termination, and recurring charges authorized to chosen payment providers. Subscription tiers appear to wrap project-based deliverables into ongoing managed-service engagements.
- Verifier & Evaluation Engineering Engagements: Specialized engagements designing custom RL post-training verifiers (rubrics, agentic graders, programmatic checks) and benchmarks (LEAPBench, AttuneBench) for frontier labs. Sold as high-value, judgment-encoding work where verification methodology is the deliverable.
Pricing tiers
| Model | Billing | Price |
|---|---|---|
| Other | Monthly | Quote-based, project-scoped engagements initiated via "Request a Project" form |
Go-to-market motion3 records
Distribution channels3 records
Marketing channels8 records
Pareto AI product offering
Product offeringCore offering
Pareto AI sources vetted domain experts (physicians, lawyers, PhDs, engineers, ML developers, finance professionals, linguists) through its proprietary Forte platform and delivers expert-graded AI training data, rubric design, and verifier engineering services to frontier AI labs, government safety institutes, and academic research consortia. Engagements are project-scoped and include harm-scale datasets, debate judgement pipelines, RL reward-signal verifiers, and open benchmarks for evaluating frontier LLMs.
Product overview
Pareto AI (Pareto, Inc.) operates as an integrated platform-plus-services business for expert human data and AI training: the public Pareto website (pareto.ai) is the corporate and marketing surface, while the core product is the Forte platform (forte.pareto.ai), an expert portal where vetted domain specialists — in Medicine & health, Law & policy, STEM & research, Finance & strategy, Software & ML, and Language & humanities — produce annotation, rubric grading, debate judgement, and verification data for frontier-lab clients. On top of Forte, Pareto delivers project-based data collection services (e.g., the Human-Judged LLM Debate Judgement Pipeline for MATS/Anthropic/UCL and the Harmful-Advice Classifier collaboration with the UK AISI) and publishes research artifacts that double as products, notably the open LEAPBench and AttuneBench benchmarks, which use Forte-style expert labelling and verifier pipelines to evaluate frontier LLMs. Forte and Pareto Projects (Beta) is the early-stage community layer that surrounds the core platform.
Differentiator
Problem solved
Functional benefit
Brands
- Forte: The branded experts platform operated by Pareto at forte.pareto.ai, used by AI trainers/experts to register, complete projects, and access work opportunities.
Products and services
- Forte Expert Platform
- Custom AI Training Data Projects
- Verifier and Evaluation Engineering Engagements
Quantifiable outcome
- Fine-tuned Llama-3.1-8B accuracy on harmful-advice detection rose from 77% to 96% (a 19 percentage point lift), beating GPT-4o zero-shot at 93%.
- +5 more outcomes
Companies that use Pareto AI
Customer profileNamed customers7 records
Ideal customer profiles3 records
Pareto AI technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
Integration10 records
AI capability12 records
Feature6 records
Pareto AI partnerships and signals
Strategic signalPartnerships
Four partnerships are on record, tiered flagship and core.
- UK AI Safety Institute (AISI)flagshipDeep collaboration: Pareto sourced licensed doctors, therapists, and career coaches to co-design 0–4 harm rubrics and grade 6,707 examples; the resulting classifier was fine-tuned into Llama-3.1-8B and deployed live in an AISI study protecting 2,302 participants. AISI open-sourced the paper, model, and dataset, and credited Pareto team members Elizabeth Nguyen and Daria Butuc alongside AISI staff Lennart Luettgau and Henry Davidson.
- ML Alignment & Theory Scholars (MATS) ProgramcoreCommissioned Pareto to collect human-judged debate data for scalable-oversight research. Pareto sourced and onboarded 20 experts in under a month via referral-based sourcing, supported iterative weekly guideline changes, and enabled direct researcher-worker communication. Outputs underpinned research by Dan Valentine, Akbir Khan, and others.
- Thoughtful LabcoreCo-developed AttuneBench with Pareto's Research team (led by Mark Whiting). Thoughtful Lab, led by Karina Nguyen, contributed dataset creation and analysis; participants supplied 50,000+ annotations across 200 conversations. Benchmark code and leaderboard are open-sourced.
- University College London (UCL)coreUCL researchers participated in the MATS-led human-judged debate study for which Pareto supplied the expert data; findings advanced the paper "Scalable AI Safety via Doubly-Efficient Debate" (arXiv 2402.06782).
Scale indicators12 records
Recent moves5 records
Expansion highlights6 records
Pareto AI competitors and assessment
Company assessmentDirect peers
- Surge AI: Surge AI is a direct competitor offering expert-labeled RLHF and RLVR data to frontier AI labs, with a similarly specialist-talent focus (often humanities-trained labelers) and enterprise sales motion. It targets the same buyer set (OpenAI, Anthropic, Google) and competes head-to-head on rubric-driven human-judgment tasks.
- Mercor: Mercor is a direct competitor that matches domain experts (engineers, scientists, lawyers, doctors) to AI training and evaluation projects, using a talent-marketplace model. It overlaps with Pareto on expert sourcing for frontier-lab post-training and on professional-domain coverage.
Broad incumbents
- Scale AI: Scale AI is the dominant incumbent in AI training data, serving the same frontier-lab customers with a much broader portfolio spanning labeling, evaluation, RLHF, government, and autonomous-vehicle data. While its core is task-based labeling rather than specialist expert sourcing, it is the default vendor Pareto must displace or complement.
- Labelbox: Labelbox provides an enterprise data-labeling platform combining software tooling with managed labeling services, including expert annotation workflows for generative-AI use cases. It is comparable on the platform-plus-services model and on serving enterprise AI teams, though its primary moat is software rather than expert sourcing.
- Appen: Appen is a long-standing publicly-traded data-annotation vendor with a large global crowd workforce, serving enterprise AI teams across NLP, vision, and speech. It is comparable as a broad incumbent in human-generated training data, though its positioning is more commodity and less specialist-expert focused than Pareto's.
- TELUS International (formerly Lionbridge AI): TELUS International's AI Data Solutions division delivers large-scale multilingual data annotation and AI training data to enterprise buyers, drawing on a global workforce. It is comparable as a broad incumbent in AI training data services, though its sweet spot is more multilingual/generalist than Pareto's expert-domain focus.
Emerging players
- Snorkel AI: Snorkel AI builds programmatic data-labeling and weak-supervision tooling for enterprise ML and LLM workflows, with an emerging focus on LLM evaluation and alignment. It is comparable on the methodology side (rubrics, programmatic labels, expert-in-the-loop) but competes more on platform tooling than on a managed expert workforce.
- Sama: Sama delivers AI data annotation and model-validation services with a focus on high-quality, ethically-sourced workforce across computer vision and NLP. It overlaps with Pareto on professional managed-services annotation for enterprise AI teams and on ethical-workforce positioning.
- CloudFactory: CloudFactory provides managed data-labeling teams for AI/ML pipelines, combining a trained cloud workforce with workflow tooling. It is comparable on the managed-services data-labeling model and on serving enterprise AI buyers, though its expert credentialing is less deep than Pareto's.
- Encord: Encord offers a data-development platform for AI teams spanning labeling, evaluation, and curation, with workflows designed for multimodal and RLHF pipelines. It is comparable on the platform-and-tooling side of expert annotation and evaluation, though it is less specialized in frontier-lab post-training than Pareto.
Market position
Strengths4 records
Weaknesses5 records
Competitive moat5 records
Key risks6 records
Key highlights6 records
Customer concentration
Pareto AI social profiles
Digital presencePareto AI compliance and trust
Trust signalCompliance4 records
Pareto AI financial estimates
Financial estimateRevenue estimate
Valuation estimate
Pareto AI leadership team
Management profileNumber of profiles
Profiles2 records
Pareto AI funding detail
Funding detailFunding overview
Funding rounds3 records
Investors9 records
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Pareto AI M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Pareto AI
What does Pareto AI do?
Pareto AI sources vetted domain experts (physicians, lawyers, PhDs, engineers, ML developers, finance professionals, linguists) through its proprietary Forte platform and delivers expert-graded AI training data, rubric design, and verifier engineering services to frontier AI labs, government safety institutes, and academic research consortia. Engagements are project-scoped and include harm-scale datasets, debate judgement pipelines, RL reward-signal verifiers, and open benchmarks for evaluating frontier LLMs.
Is Pareto AI a public or private company?
Pareto AI is a private company. It is classified as venture growth investor backed and is currently operating.
When was Pareto AI founded?
Pareto AI was founded in 2020. It employs 101 to 250 people.
Where is Pareto AI based?
Pareto AI is headquartered in Wilmington, United States, in the North America region.
How does Pareto AI make money?
Three revenue lines are on record. Custom AI Training Data Projects are the primary driver. The others are subscription / Recurring Service Plans and verifier & Evaluation Engineering Engagements.
Who are Pareto AI's main competitors?
Direct peers on record are Surge AI and Mercor. Broad incumbents are Scale AI, Labelbox, Appen and TELUS International (formerly Lionbridge AI). Emerging players are Snorkel AI, Sama, CloudFactory and Encord.
Does Pareto AI have an API?
No public API is recorded for Pareto AI.
What industry is Pareto AI in?
Pareto AI's product category is Expert AI Training Data Platform. Its primary akta.pro industry code is HDAAALAK, Bias, Fairness & Representativeness Testing for Datasets, with a secondary code of HDAAAMAL, Safety & Alignment Evaluation (red-teaming, harmful capability testing). Its NAICS code is 5417 and its SIC code is 7371.