Metaphi
Metaphi builds reinforcement learning environments and specialized benchmarks (COBOLBench, FigmaBench, VideoBench) for frontier AI companies to train and evaluate autonomous coding and generation agents on enterprise-grade tasks.
- Company typePrivate
- Founded2022
- HeadquartersMckinney, United States
- Headcount1–10
- GTM typeB2B
- OfferingSoftware
What Metaphi does
Metaphi is an AI infrastructure company that builds reinforcement learning (RL) environments and benchmarks for frontier AI companies developing autonomous coding and generation agents. The company's core product, Simhub, provides interactive simulation environments that move agent evaluation beyond static, non-interactive benchmarks, drawing on principles from autonomous vehicle simulation. Built on this platform are specialized public benchmarks including COBOLBench (100 enterprise COBOL maintenance tasks), FigmaBench (production Figma-to-code conversion), VideoBench (183 chapters across 41 courses for video/animation/presentation generation), and the LegacySWE research initiative on legacy enterprise system modernization. These products are delivered through the evals.metaphi.ai evaluation platform, which hosts a long-horizon agent leaderboard and targets frontier AI labs seeking realistic RL training and evaluation environments for out-of-distribution enterprise domains.
The company operates a research-first go-to-market motion, publishing technical essays and benchmarks as top-of-funnel demand generation, with enterprise sales conducted via direct contact forms capturing domain interest (Enterprise, Coding, Physical AI). Revenue is described as subscription-recurring from custom RL training environment contracts with frontier AI companies; pricing is not publicly disclosed and is quote-based. Founded in 2022 and based in McKinney, Texas, Metaphi has 1-10 employees, is led by co-founder Abhishek Chandwani, and is currently hiring RL engineers and ex-founders to extend into Physical AI and other unexplored domains. The company has raised $1.25M in seed funding from Flybridge (July 2022) with no subsequent rounds disclosed in available data.
Metaphi firmographics
Firmographics- Name
- Metaphi
- Website
- https://metaphi.ai
- Company type
- Private
- Founded year
- 2022
- Operating status
- Operating
- Headcount range
- 1–10 employees
- Short description
- Metaphi builds reinforcement learning environments and specialized benchmarks (COBOLBench, FigmaBench, VideoBench) for frontier AI companies to train and evaluate autonomous coding and generation agents on enterprise-grade tasks.
- Ownership category
- akta.pro rank
Metaphi industry classification
Industry- Product category
- AI Evaluation and RL Training Infrastructure
- NAICS
- Custom Computer Programming Services (541511)
- SIC
- Services-Prepackaged Software (7372)
- akta.pro primary industry
- RLHF, Human Feedback & Evaluation Data Tools (HDAAALAJ)
- akta.pro secondary industry
- End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management) (HDAEANAA)
Keywords
Where Metaphi is headquartered
LocationHeadquarters
- HQ city
- Mckinney
- HQ country
- United States
- HQ region
- North America
Markets served
Metaphi business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Technology or R&D, Personnel, Marketing or Sales
Revenue model
- Enterprise RL Environments: Custom RL training environments built for frontier AI companies addressing out-of-distribution enterprise domains. Revenue derived from enterprise contracts for evaluation and training infrastructure.
Go-to-market motion1 record
Distribution channels1 record
Marketing channels3 records
Metaphi product offering
Product offeringCore offering
Metaphi builds reinforcement learning (RL) environments and evaluation benchmarks for frontier AI companies developing coding agents. The core offering is delivered through the evals.metaphi.ai platform, which provides interactive simulation environments (Simhub) and domain-specific benchmarks including COBOLBench (enterprise COBOL maintenance), FigmaBench (Figma-to-code conversion), and VideoBench (video/animation generation). These tools enable autonomous improvement of AI agents on out-of-distribution enterprise tasks beyond static, non-interactive benchmarks.
Product overview
Metaphi builds RL (reinforcement learning) environments for frontier AI companies, offering a platform architecture centered on Simhub for interactive simulation environments. The core product is delivered through evals.metaphi.ai with two tiers: Enterprise for out-of-distribution enterprise domain evaluation, and Long-Horizon for extended task-sequence benchmarking with leaderboards. Specialized benchmarks include COBOLBench (100-task COBOL maintenance benchmark), VideoBench (video/animation generation from data rooms), and FigmaBench (Figma-to-code conversion). LegacySWE is a research initiative focused on legacy enterprise system maintenance. The platform tests AI coding agents through API interactions, design hierarchy extraction, and iterative deployment across enterprise domains including Coding, Enterprise, and Physical AI.
Differentiator
Problem solved
Functional benefit
Brands
- Simhub: Interactive simulation environments that move agent evaluation beyond static, non-interactive benchmarks.
- COBOLBench
- FigmaBench
- VideoBench
- LegacySWE
Products and services
- Simhub Interactive simulation environments that move AI agent evaluation beyond static, non-interactive benchmarks. Built on principles from autonomous vehicle simulation to transform code generation agent training and evaluation for frontier AI companies.
- COBOLBench A 100-task public benchmark drawn from real enterprise COBOL maintenance environments, measuring frontier coding systems on legacy enterprise COBOL maintenance and modernization tasks. For frontier AI companies developing coding agents.
- VideoBench A benchmark measuring AI agents on video, animation, and presentation generation from curated data rooms, covering 183 chapters across 41 courses. For frontier AI companies developing video/animation generation agents.
- FigmaBench A benchmark measuring AI agents on production Figma-to-code conversion through API interaction, design hierarchy extraction, and iterative deployment. For frontier AI companies developing design-to-code conversion agents.
- Metaphi Evaluations Platform (evals.metaphi.ai) Enterprise AI evaluation platform providing RL environments and long-horizon agent leaderboards for measuring AI coding agents on out-of-distribution enterprise domains. Combines the Enterprise evaluation tier and Long-Horizon leaderboard tier as a single integrated platform.
Companies that use Metaphi
Customer profileSegments2 records
Ideal customer profiles1 record
Metaphi technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability7 records
Feature5 records
Metaphi partnerships and signals
Strategic signalScale indicators2 records
Recent moves6 records
Expansion highlights5 records
Metaphi competitors and assessment
Company assessmentDirect peers
- Braintrust: Provides an LLM evaluation, observability, and prompt engineering platform for AI teams building production applications. Directly comparable to Metaphi's enterprise evaluation platform (evals.metaphi.ai) and leaderboard-based approach to benchmarking AI agent quality.
- LangSmith (LangChain): Offers evaluation, tracing, and monitoring tooling for LLM-powered applications and agents. Comparable to Metaphi in serving AI engineering teams that need structured evaluation infrastructure beyond static benchmarks.
- Galileo: Builds an evaluation and observability platform specifically for LLM and generative AI applications. Overlaps with Metaphi's focus on agent evaluation, especially for production-grade AI systems.
- HumanLoop: Provides an LLM evaluation and experimentation platform aimed at teams shipping AI products. Comparable as a vendor helping enterprises evaluate and iterate on generative AI agents.
Broad incumbents
- Arize AI: ML/LLM observability and evaluation platform serving production AI teams. Broader than Metaphi (covers traditional ML observability too), but directly competes in LLM/agent evaluation workflows.
- Scale AI: Major RLHF, data labeling, and evaluation vendor serving frontier AI labs. Much larger and broader than Metaphi, but overlapping in supplying evaluation data and benchmark infrastructure to frontier model developers.
- Surge AI: Data labeling and RLHF provider for frontier AI companies. Compares to Metaphi as a supplier of training/evaluation data infrastructure, though Surge is broader across modalities.
- Weights & Biases: Established ML and AI development platform with evaluation, experiment tracking, and model registry capabilities. A broader incumbent whose eval tooling overlaps with Metaphi's agent evaluation focus.
Emerging players
- SWE-bench: Open-source benchmark for evaluating AI systems on real-world software engineering tasks. Directly comparable to Metaphi's COBOLBench and LegacySWE efforts, though focused on Python repos rather than legacy enterprise code.
- Confident AI (DeepEval): Open-source LLM evaluation framework used by AI engineers to benchmark and test generative AI applications. Comparable to Metaphi as a vendor targeting teams that need structured agent/LLM evaluation.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat4 records
Key risks6 records
Key highlights6 records
Customer concentration
Metaphi financial estimates
Financial estimateRevenue estimate
Valuation estimate
Metaphi leadership team
Management profileNumber of profiles
Profiles1 record
Metaphi funding detail
Funding detailFunding overview
Funding rounds1 record
Investors1 record
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Metaphi M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Metaphi
What does Metaphi do?
Metaphi builds reinforcement learning (RL) environments and evaluation benchmarks for frontier AI companies developing coding agents. The core offering is delivered through the evals.metaphi.ai platform, which provides interactive simulation environments (Simhub) and domain-specific benchmarks including COBOLBench (enterprise COBOL maintenance), FigmaBench (Figma-to-code conversion), and VideoBench (video/animation generation). These tools enable autonomous improvement of AI agents on out-of-distribution enterprise tasks beyond static, non-interactive benchmarks.
When was Metaphi founded?
Metaphi was founded in 2022. It employs 1 to 10 people.
Where is Metaphi based?
Metaphi is headquartered in Mckinney, United States, in the North America region.
How does Metaphi make money?
One revenue line is on record: enterprise RL Environments.
Who are Metaphi's main competitors?
Direct peers on record are Braintrust, LangSmith (LangChain), Galileo and HumanLoop. Broad incumbents are Arize AI, Scale AI, Surge AI and Weights & Biases. Emerging players are SWE-bench and Confident AI (DeepEval).
Does Metaphi have an API?
No public API is recorded for Metaphi.
What industry is Metaphi in?
Metaphi's product category is AI Evaluation and RL Training Infrastructure. Its primary akta.pro industry code is HDAAALAJ, RLHF, Human Feedback & Evaluation Data Tools, with a secondary code of HDAEANAA, End-to-End Enterprise AI Platforms (MLOps & Model Lifecycle Management). Its NAICS code is 541511 and its SIC code is 7372.