Model Evaluation & Threat Research
METR is a 501(c)(3) research nonprofit that builds scientific benchmarks and evaluation methodologies to measure autonomous capabilities and catastrophic risks of frontier AI models, serving AI developers, governments, and the research community without charging for its work.
- Company typePrivate
- Founded2023
- HeadquartersBerkeley, United States
- Headcount11–50
- GTM typeB2B
- OfferingSoftware
What Model Evaluation & Threat Research does
Model Evaluation & Threat Research (METR) is a 501(c)(3) research nonprofit incorporated in Delaware (EIN 99-1219864), operating from Covina, California. Founded around 2022-2023 as ARC Evals and rebranded to METR, the organization develops scientific methodologies to measure whether and when AI systems might threaten catastrophic harm, with a specific focus on autonomous capabilities and AI R&D acceleration of frontier AI models.
METR's core technical stack comprises the Time Horizon metric (measuring the length of software tasks AI agents can complete, currently at version 1.1 with a documented ~7-month doubling time over six years), the open-source Hawk evaluation platform built on Inspect AI, the Vivaria evaluation infrastructure, and the METR Task Standard for portable evaluation tasks. Supporting research assets include the HCAST (Human-Calibrated Autonomy Software Tasks) benchmark, RE-Bench for ML research engineering tasks, MirrorCode for weeks-long coding tasks, and the MALT dataset cataloguing evaluation integrity threats such as reward hacking and sandbagging. The organization also publishes Frontier Risk Reports and a Frontier AI Safety Policies resource hub covering 12 frontier AI companies.
METR serves three primary constituencies: frontier AI developers (Anthropic, OpenAI, Google DeepMind, Meta, xAI), governments and regulators (US AI Safety Institute, EU AI Office, EU AI Act, California SB 53, NY RAISE Act), and the AI research community. The organization explicitly does not accept compensation for evaluation work; instead it is funded through donations and grants, with frontier AI companies providing model access and compute credits in kind rather than cash. GTM is content- and research-driven via open publications, open-source GitHub releases, earned media coverage in major technology publications, and direct advisory relationships.
Model Evaluation & Threat Research firmographics
Firmographics- Name
- Model Evaluation & Threat Research
- Legal name
- Model Evaluation and Threat Research, Inc.
- Website
- https://metr.org
- Company type
- Private
- Founded year
- 2023
- Operating status
- Operating
- Headcount range
- 11–50 employees
- Short description
- METR is a 501(c)(3) research nonprofit that builds scientific benchmarks and evaluation methodologies to measure autonomous capabilities and catastrophic risks of frontier AI models, serving AI developers, governments, and the research community without charging for its work.
- Ownership category
- akta.pro rank
Model Evaluation & Threat Research industry classification
Industry- Product category
- AI Safety Evaluation Software
- NAICS
- Scientific Research and Development Services (5417), Other Scientific and Technical Consulting Services (541690)
- SIC
- Services-Management Consulting Services (8742), Services-Commercial Physical & Biological Research (8731)
- akta.pro primary industry
- Safety & Alignment Evaluation (red-teaming, harmful capability testing) (HDAAAMAL)
- akta.pro secondary industries
- Enterprise AI Governance, Risk & Compliance Platforms (Model Risk, Audit, Policies) (HDAEANAE), Model Governance, Risk & Compliance (GRC) Platforms (HDAAAKAA)
Keywords
Where Model Evaluation & Threat Research is headquartered
LocationHeadquarters
- HQ city
- Berkeley
- HQ country
- United States
- HQ region
- North America
Offices1 record
Markets served
Model Evaluation & Threat Research business model
Business model- GTM type
- B2B
- Offering type
- Software
- Cost components
- Personnel, Technology or R&D, Operations
Revenue model
- Research Grants and Donations: METR is a 501(c)(3) nonprofit funded by donations and grants. The organization does not accept compensation for its evaluation work.
Distribution channels3 records
Marketing channels7 records
Model Evaluation & Threat Research product offering
Product offeringCore offering
METR is a 501(c)(3) nonprofit research organization that develops methodologies, benchmarks, datasets, and software platforms to scientifically evaluate frontier AI systems' autonomous capabilities and catastrophic risks. Its core offerings include the Time Horizon metric for measuring AI task completion lengths, open-source evaluation platforms (Hawk, Vivaria), and benchmarks (RE-Bench, HCAST, MirrorCode) that quantify how long and complex software tasks AI agents can complete autonomously. METR also publishes the Frontier Risk Report and provides independent advisory services to AI developers and governments on risk assessment methodologies.
Product overview
METR (Model Evaluation & Threat Research) is a research nonprofit that offers a suite of AI evaluation tools and benchmarks focused on assessing frontier AI capabilities and risks. The core offerings include the Time Horizon metric for measuring AI task completion capabilities, the Hawk platform for running evaluations at scale, and specialized benchmarks like RE-Bench (ML research tasks), MirrorCode (weeks-long coding tasks), and HCAST (autonomy software tasks). Supporting resources include the MALT dataset for evaluation integrity threats, Vivaria evaluation infrastructure, and the Frontier AI Safety Policies hub. These products work together to provide comprehensive assessment of whether and when AI systems might threaten catastrophic harm, with particular focus on autonomous capabilities and AI R&D acceleration.
Differentiator
Problem solved
Functional benefit
Products and services
- Time Horizon Metric A methodology for measuring AI performance based on the length of software tasks AI agents can autonomously complete, demonstrating exponential improvement over time. For use by AI researchers, developers, and policymakers to assess frontier AI autonomy.
- Hawk An open-source platform for running AI agent evaluations at scale, built upon the Inspect AI framework. For AI researchers and developers conducting capability evaluations.
- Vivaria METR's tool for running evaluations and conducting agent elicitation research. For AI researchers performing large-scale agent evaluations.
- MALT (Manually-reviewed Agentic Labeled Transcripts) A dataset of natural and prompted examples of behaviors that threaten evaluation integrity, including generalized reward hacking and sandbagging. For AI safety researchers building more robust evaluation methods.
- HCAST (Human-Calibrated Autonomy Software Tasks) A benchmark for measuring the abilities of frontier AI systems to complete diverse software tasks autonomously. For AI researchers measuring autonomy capabilities.
- RE-Bench A benchmark for measuring the performance of humans and frontier model agents on ML research engineering tasks. For AI researchers and developers evaluating R&D capabilities.
- MirrorCode A benchmark providing evidence that AI can already complete weeks-long coding tasks, including reimplementing a 16,000-line codebase. For AI researchers assessing long-horizon autonomous coding capabilities.
- METR Task Standard A published standard for defining tasks for evaluating the capabilities of AI agents, enabling portable evaluation tasks. For AI researchers and developers building evaluation infrastructure.
Quantifiable outcome
- Time horizon doubling time of ~7 months for frontier AI models
- +2 more outcomes
Companies that use Model Evaluation & Threat Research
Customer profileSegments3 records
Ideal customer profiles3 records
Model Evaluation & Threat Research technology and API
TechnologyTechnology focussed Yes
API detail
- Has API
- No
- API docs
- API detail
Core technology
AI maturity
App detail
AI capability8 records
Feature6 records
Model Evaluation & Threat Research partnerships and signals
Strategic signalPartnerships
Five partnerships are on record, tiered core and minor.
- AnthropiccoreMETR conducts independent evaluations of Anthropic models including Claude 3.5, 3.7, and Opus 4.6. Anthropic provides non-public information and model access for evaluation research. METR also conducted red-teaming of Anthropic's internal agent monitoring systems and reviewed Anthropic's Risk Reports including the sabotage risk and automated R&D sections.
- Google DeepMindcoreGoogle DeepMind participated in the Frontier Risk Report pilot exercise and provides model access for METR evaluations. METR has conducted evaluations of Google models and the company has contributed to frontier AI safety discussions.
- MetacoreMeta participated in the Frontier Risk Report pilot exercise alongside Anthropic, Google, and OpenAI. METR evaluates Meta's frontier AI models and provides independent assessment of AI capabilities and risks.
- OpenAIcoreOpenAI participated in the Frontier Risk Report pilot exercise and has provided model access and compute credits for METR's evaluation research. METR has evaluated GPT-4o, o1-preview, o3, GPT-5, and GPT-5.1-Codex-Max models.
- xAIminorxAI has provided access and compute credits to support METR's evaluation research, though their involvement in specific evaluations appears less extensive than other frontier AI labs.
Scale indicators4 records
Recent moves3 records
Expansion highlights5 records
Model Evaluation & Threat Research competitors and assessment
Company assessmentDirect peers
- Center for AI Safety (CAIS): Nonprofit research organization focused on catastrophic AI risks, including safety evaluations, technical research, and policy advocacy. Operates in the same AI safety evaluation and policy space as METR.
- Apollo Research: Independent AI safety research organization focused on evaluating deceptive and scheming behaviors in frontier models. Closest direct competitor to METR, working with similar labs on pre-deployment capability evaluations.
- Redwood Research: Nonprofit AI safety research organization focused on alignment and deception. Operates in adjacent AI safety research space with similar nonprofit structure and elite research talent base.
- Future of Life Institute: Nonprofit focused on AI safety, governance, and policy advocacy. Comparable mission and funding model (donations), with overlap in regulatory advisory and catastrophic risk framing.
Broad incumbents
- MITRE Corporation: Federally-funded research organization that operates AI safety and assurance initiatives. Comparable as an independent evaluator serving government and industry, but with vastly broader portfolio.
- US AI Safety Institute (CAISI): US government body within NIST conducting frontier model evaluations. METR has advisory relationships with this body; operates as a government-mandated evaluation counterpart to METR's independent nonprofit role.
- UK AI Security Institute (formerly UK AI Safety Institute): Government-backed AI evaluation body conducting pre-deployment assessments of frontier models. Comparable in mission to METR's evaluations but government-funded and operating at larger scale across broader mandate.
Emerging players
- Conjecture: AI alignment and safety company focused on controlling advanced AI systems. Operates in adjacent AI safety research space with commercial orientation compared to METR's nonprofit model.
- SaferAI: Nonprofit focused on AI governance and risk assessment for frontier AI developers. Overlaps with METR's Frontier AI Safety Policies hub and policy advisory work, though narrower technical scope.
Others
- Anthropic Responsible Scaling Policy Team: Internal evaluation and policy team at Anthropic that METR independently reviews. Functions as both a customer/partner relationship and a competing in-house evaluation capability that could reduce reliance on external evaluators like METR.
Market position
Strengths4 records
Weaknesses4 records
Competitive moat5 records
Key risks5 records
Key highlights6 records
Customer concentration
Model Evaluation & Threat Research social profiles
Digital presenceModel Evaluation & Threat Research financial estimates
Financial estimateRevenue estimate
Valuation estimate
Model Evaluation & Threat Research leadership team
Management profileNumber of profiles
Profiles2 records
Model Evaluation & Threat Research funding detail
Funding detailFunding overview
Funding rounds
Investors
Funding detail is available on the Subscription and Enterprise plan.Contact sales →
Model Evaluation & Threat Research M&A and investment
M&A and investmentM&A
Investments
M&A and investment is available on the Subscription and Enterprise plan.Contact sales →
Frequently asked questions about Model Evaluation & Threat Research
What does Model Evaluation & Threat Research do?
METR is a 501(c)(3) nonprofit research organization that develops methodologies, benchmarks, datasets, and software platforms to scientifically evaluate frontier AI systems' autonomous capabilities and catastrophic risks. Its core offerings include the Time Horizon metric for measuring AI task completion lengths, open-source evaluation platforms (Hawk, Vivaria), and benchmarks (RE-Bench, HCAST, MirrorCode) that quantify how long and complex software tasks AI agents can complete autonomously. METR also publishes the Frontier Risk Report and provides independent advisory services to AI developers and governments on risk assessment methodologies.
Is Model Evaluation & Threat Research a public or private company?
Model Evaluation & Threat Research is a private company. It is classified as nonprofit foundation owned and is currently operating.
When was Model Evaluation & Threat Research founded?
Model Evaluation & Threat Research was founded in 2023. It employs 11 to 50 people.
Where is Model Evaluation & Threat Research based?
Model Evaluation & Threat Research is headquartered in Berkeley, United States, in the North America region.
How does Model Evaluation & Threat Research make money?
One revenue line is on record: research Grants and Donations.
Who are Model Evaluation & Threat Research's main competitors?
Direct peers on record are Center for AI Safety (CAIS), Apollo Research, Redwood Research and Future of Life Institute. Broad incumbents are MITRE Corporation, US AI Safety Institute (CAISI) and UK AI Security Institute (formerly UK AI Safety Institute). Emerging players are Conjecture and SaferAI. Anthropic Responsible Scaling Policy Team is listed as an others.
Does Model Evaluation & Threat Research have an API?
No public API is recorded for Model Evaluation & Threat Research.
What industry is Model Evaluation & Threat Research in?
Model Evaluation & Threat Research's product category is AI Safety Evaluation Software. Its primary akta.pro industry code is HDAAAMAL, Safety & Alignment Evaluation (red-teaming, harmful capability testing), with a secondary code of HDAEANAE, Enterprise AI Governance, Risk & Compliance Platforms (Model Risk, Audit, Policies). Its NAICS code is 5417 and its SIC code is 8742.