Developer docs
API playgroundTry for free, no card

Search company profiles

Model Evaluation & Threat Research

Full company profile

uuid006fqux

Namestring
Model Evaluation & Threat Research
Legal namestring
Model Evaluation and Threat Research, Inc.
Websiteurl
metr.org
Company typeenum
Private
Founded yearint
2023
Descriptiontext

Model Evaluation & Threat Research (METR) is a 501(c)(3) research nonprofit incorporated in Delaware (EIN 99-1219864), operating from Covina, California. Founded around 2022-2023 as ARC Evals and rebranded to METR, the organization develops scientific methodologies to measure whether and when AI systems might threaten catastrophic harm, with a specific focus on autonomous capabilities and AI R&D acceleration of frontier AI models.

METR's core technical stack comprises the Time Horizon metric (measuring the length of software tasks AI agents can complete, currently at version 1.1 with a documented ~7-month doubling time over six years), the open-source Hawk evaluation platform built on Inspect AI, the Vivaria evaluation infrastructure, and the METR Task Standard for portable evaluation tasks. Supporting research assets include the HCAST (Human-Calibrated Autonomy Software Tasks) benchmark, RE-Bench for ML research engineering tasks, MirrorCode for weeks-long coding tasks, and the MALT dataset cataloguing evaluation integrity threats such as reward hacking and sandbagging. The organization also publishes Frontier Risk Reports and a Frontier AI Safety Policies resource hub covering 12 frontier AI companies.

METR serves three primary constituencies: frontier AI developers (Anthropic, OpenAI, Google DeepMind, Meta, xAI), governments and regulators (US AI Safety Institute, EU AI Office, EU AI Act, California SB 53, NY RAISE Act), and the AI research community. The organization explicitly does not accept compensation for evaluation work; instead it is funded through donations and grants, with frontier AI companies providing model access and compute credits in kind rather than cash. GTM is content- and research-driven via open publications, open-source GitHub releases, earned media coverage in major technology publications, and direct advisory relationships.

Short descriptiontext

METR is a 501(c)(3) research nonprofit that builds scientific benchmarks and evaluation methodologies to measure autonomous capabilities and catastrophic risks of frontier AI models, serving AI developers, governments, and the research community without charging for its work.

Operating statusenum
Operating
Ownership categoryenum
Headcount rangeband
11–50
akta.pro rankint
HeadquartersBerkeley, United States
HQ citystring
Berkeley
HQ countrystring
United States
HQ regionstring
North America
Markets served

Serves global market

Offices1 record

Each record includes

City, Country, Type, Description, Source

Keyword5 values
AI safety evaluation, frontier AI assessment, AI capability benchmarks, autonomous AI testing, AI risk research
Industry3 codes
1Safety & Alignment Evaluation (red-teaming, harmful capability testing)
CodeHDAAAMALPrimaryYes
2Enterprise AI Governance, Risk & Compliance Platforms (Model Risk, Audit, Policies)
CodeHDAEANAEPrimaryNo
3Model Governance, Risk & Compliance (GRC) Platforms
CodeHDAAAKAAPrimaryNo
NAICS code2 codes
  • Scientific Research and Development Services5417
  • Other Scientific and Technical Consulting Services541690
SIC code2 codes
  • Services-Management Consulting Services8742
  • Services-Commercial Physical & Biological Research8731
Product category
AI Safety Evaluation Software
Social media profiles3 records
Revenue model1 record
1Research Grants and Donations
TypeOthers
Description

METR is a 501(c)(3) nonprofit funded by donations and grants. The organization does not accept compensation for its evaluation work.

metr.org
Marketing channels7 records

Each record includes

Title, Type, Stage, Description, Source

Distribution channels3 records

Each record includes

Title, Type, Scope, Target buyer, Description, Source

Cost components3 values
Personnel, Technology or R&D, Operations
GTM typeB2B
B2B
Offering typeSoftware
Software
Core offering1 text field

METR is a 501(c)(3) nonprofit research organization that develops methodologies, benchmarks, datasets, and software platforms to scientifically evaluate frontier AI systems' autonomous capabilities and catastrophic risks. Its core offerings include the Time Horizon metric for measuring AI task completion lengths, open-source evaluation platforms (Hawk, Vivaria), and benchmarks (RE-Bench, HCAST, MirrorCode) that quantify how long and complex software tasks AI agents can complete autonomously. METR also publishes the Frontier Risk Report and provides independent advisory services to AI developers and governments on risk assessment methodologies.

Differentiator
Functional benefit
Problem solved
Quantifiable outcome1 of 3 values shown
  • Time horizon doubling time of ~7 months for frontier AI models
+2 more records
Product overview1 text field

METR (Model Evaluation & Threat Research) is a research nonprofit that offers a suite of AI evaluation tools and benchmarks focused on assessing frontier AI capabilities and risks. The core offerings include the Time Horizon metric for measuring AI task completion capabilities, the Hawk platform for running evaluations at scale, and specialized benchmarks like RE-Bench (ML research tasks), MirrorCode (weeks-long coding tasks), and HCAST (autonomy software tasks). Supporting resources include the MALT dataset for evaluation integrity threats, Vivaria evaluation infrastructure, and the Frontier AI Safety Policies hub. These products work together to provide comprehensive assessment of whether and when AI systems might threaten catastrophic harm, with particular focus on autonomous capabilities and AI R&D acceleration.

Product and service8 records
1Time Horizon Metric
CategoryEvaluation Methodology
Description

A methodology for measuring AI performance based on the length of software tasks AI agents can autonomously complete, demonstrating exponential improvement over time. For use by AI researchers, developers, and policymakers to assess frontier AI autonomy.

2Hawk
CategoryEvaluation Platform
Description

An open-source platform for running AI agent evaluations at scale, built upon the Inspect AI framework. For AI researchers and developers conducting capability evaluations.

3Vivaria
CategoryEvaluation Platform
Description

METR's tool for running evaluations and conducting agent elicitation research. For AI researchers performing large-scale agent evaluations.

4MALT (Manually-reviewed Agentic Labeled Transcripts)
CategoryDataset
Description

A dataset of natural and prompted examples of behaviors that threaten evaluation integrity, including generalized reward hacking and sandbagging. For AI safety researchers building more robust evaluation methods.

5HCAST (Human-Calibrated Autonomy Software Tasks)
CategoryBenchmark
Description

A benchmark for measuring the abilities of frontier AI systems to complete diverse software tasks autonomously. For AI researchers measuring autonomy capabilities.

6RE-Bench
CategoryBenchmark
Description

A benchmark for measuring the performance of humans and frontier model agents on ML research engineering tasks. For AI researchers and developers evaluating R&D capabilities.

7MirrorCode
CategoryBenchmark
Description

A benchmark providing evidence that AI can already complete weeks-long coding tasks, including reimplementing a 16,000-line codebase. For AI researchers assessing long-horizon autonomous coding capabilities.

8METR Task Standard
CategoryStandard
Description

A published standard for defining tasks for evaluating the capabilities of AI agents, enabling portable evaluation tasks. For AI researchers and developers building evaluation infrastructure.

Scale indicator4 records

Each record includes

Type, Value, Description, Source

Partnership5 partners
Strategic tierCoreTypeStrategic or Co-development PartnerAnnounced on2026-05-19
Description

METR conducts independent evaluations of Anthropic models including Claude 3.5, 3.7, and Opus 4.6. Anthropic provides non-public information and model access for evaluation research. METR also conducted red-teaming of Anthropic's internal agent monitoring systems and reviewed Anthropic's Risk Reports including the sabotage risk and automated R&D sections.

Strategic tierCoreTypeStrategic or Co-development PartnerAnnounced on2026-05-19
Description

Google DeepMind participated in the Frontier Risk Report pilot exercise and provides model access for METR evaluations. METR has conducted evaluations of Google models and the company has contributed to frontier AI safety discussions.

Strategic tierCoreTypeStrategic or Co-development PartnerAnnounced on2026-05-19
Description

Meta participated in the Frontier Risk Report pilot exercise alongside Anthropic, Google, and OpenAI. METR evaluates Meta's frontier AI models and provides independent assessment of AI capabilities and risks.

Strategic tierCoreTypeStrategic or Co-development PartnerAnnounced on2026-05-19
Description

OpenAI participated in the Frontier Risk Report pilot exercise and has provided model access and compute credits for METR's evaluation research. METR has evaluated GPT-4o, o1-preview, o3, GPT-5, and GPT-5.1-Codex-Max models.

Strategic tierMinorTypeStrategic or Co-development PartnerAnnounced on2026-05-19
Description

xAI has provided access and compute credits to support METR's evaluation research, though their involvement in specific evaluations appears less extensive than other frontier AI labs.

Recent move3 records

Each record includes

Date, Type, Title, Description, Source

Expansion highlight5 records

Each record includes

Type, Description

Peers10 records
TypeDirect peer
Description

Nonprofit research organization focused on catastrophic AI risks, including safety evaluations, technical research, and policy advocacy. Operates in the same AI safety evaluation and policy space as METR.

TypeBroad incumbent
Description

Federally-funded research organization that operates AI safety and assurance initiatives. Comparable as an independent evaluator serving government and industry, but with vastly broader portfolio.

TypeDirect peer
Description

Independent AI safety research organization focused on evaluating deceptive and scheming behaviors in frontier models. Closest direct competitor to METR, working with similar labs on pre-deployment capability evaluations.

TypeDirect peer
Description

Nonprofit AI safety research organization focused on alignment and deception. Operates in adjacent AI safety research space with similar nonprofit structure and elite research talent base.

5US AI Safety Institute (CAISI)
TypeBroad incumbent
Description

US government body within NIST conducting frontier model evaluations. METR has advisory relationships with this body; operates as a government-mandated evaluation counterpart to METR's independent nonprofit role.

TypeEmerging player
Description

AI alignment and safety company focused on controlling advanced AI systems. Operates in adjacent AI safety research space with commercial orientation compared to METR's nonprofit model.

TypeDirect peer
Description

Nonprofit focused on AI safety, governance, and policy advocacy. Comparable mission and funding model (donations), with overlap in regulatory advisory and catastrophic risk framing.

TypeEmerging player
Description

Nonprofit focused on AI governance and risk assessment for frontier AI developers. Overlaps with METR's Frontier AI Safety Policies hub and policy advisory work, though narrower technical scope.

TypeOthers
Description

Internal evaluation and policy team at Anthropic that METR independently reviews. Functions as both a customer/partner relationship and a competing in-house evaluation capability that could reduce reliance on external evaluators like METR.

TypeBroad incumbent
Description

Government-backed AI evaluation body conducting pre-deployment assessments of frontier models. Comparable in mission to METR's evaluations but government-funded and operating at larger scale across broader mandate.

Market position
Strengths4 records

Each record includes

Headline, Details, Source

Weaknesses4 records

Each record includes

Headline, Details, Source

Competitive moat5 records

Each record includes

Type, Details

Key risks5 records

Each record includes

Headline, Details, Source

Key highlights6 records

Each record includes

Headline, Details, Source

Customer concentration

Classification, Details

Segment3 records

Each record includes

Title, Type, Primary, Description, Pain point addressed, Use case, Source

Ideal customer profile3 records

Each record includes

Profile, Firmographic size, Sales motion, Sales cycle length, Buying structure, Purchase trigger, Buyer persona, Geography, Industry vertical, Primary use case, Description, Pain points, Evidence proof points, Target buyer

Technology focused
Yes
API detail
Has APIbool
No

Docs URL, Description

AI capability8 records

Each record includes

Type, Description, Source

AI maturity
App detail

Has app

Feature6 records

Each record includes

Title, Differentiator, Description, Source

Core technology
Revenue estimate
Valuation estimate
Number of profiles
Profiles2 records

Each record includes

Name, Designation, Designation category, Overview, Profile commentary, Source

No data
No data
Funding overview

Funding stage, Last funding date, Total funding USD

Funding rounds

Each record includes

Round, Amount USD, Date, Pre money valuation, Total investors, Investors, News

Investors

Each record includes

Name, Type, Date of entry, Rounds participated, Website

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

M&A

Each record includes

Name, Acquisition type, Announced date, Completed date, Status, Website, News

Investment

Each record includes

Name, Round, Announced date, Lead investor, Website, News

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Model Evaluation & Threat Research

AI Safety Evaluation Softwaremetr.org

METR is a 501(c)(3) research nonprofit that builds scientific benchmarks and evaluation methodologies to measure autonomous capabilities and catastrophic risks of frontier AI models, serving AI developers, governments, and the research community without charging for its work.

What Model Evaluation & Threat Research does

Model Evaluation & Threat Research (METR) is a 501(c)(3) research nonprofit incorporated in Delaware (EIN 99-1219864), operating from Covina, California. Founded around 2022-2023 as ARC Evals and rebranded to METR, the organization develops scientific methodologies to measure whether and when AI systems might threaten catastrophic harm, with a specific focus on autonomous capabilities and AI R&D acceleration of frontier AI models.

METR's core technical stack comprises the Time Horizon metric (measuring the length of software tasks AI agents can complete, currently at version 1.1 with a documented ~7-month doubling time over six years), the open-source Hawk evaluation platform built on Inspect AI, the Vivaria evaluation infrastructure, and the METR Task Standard for portable evaluation tasks. Supporting research assets include the HCAST (Human-Calibrated Autonomy Software Tasks) benchmark, RE-Bench for ML research engineering tasks, MirrorCode for weeks-long coding tasks, and the MALT dataset cataloguing evaluation integrity threats such as reward hacking and sandbagging. The organization also publishes Frontier Risk Reports and a Frontier AI Safety Policies resource hub covering 12 frontier AI companies.

METR serves three primary constituencies: frontier AI developers (Anthropic, OpenAI, Google DeepMind, Meta, xAI), governments and regulators (US AI Safety Institute, EU AI Office, EU AI Act, California SB 53, NY RAISE Act), and the AI research community. The organization explicitly does not accept compensation for evaluation work; instead it is funded through donations and grants, with frontier AI companies providing model access and compute credits in kind rather than cash. GTM is content- and research-driven via open publications, open-source GitHub releases, earned media coverage in major technology publications, and direct advisory relationships.

Model Evaluation & Threat Research firmographics

Firmographics
Name
Model Evaluation & Threat Research
Legal name
Model Evaluation and Threat Research, Inc.
Website
https://metr.org
Company type
Private
Founded year
2023
Operating status
Operating
Headcount range
11–50 employees
Short description
METR is a 501(c)(3) research nonprofit that builds scientific benchmarks and evaluation methodologies to measure autonomous capabilities and catastrophic risks of frontier AI models, serving AI developers, governments, and the research community without charging for its work.
Ownership category
akta.pro rank

Model Evaluation & Threat Research industry classification

Industry
Product category
AI Safety Evaluation Software
NAICS
Scientific Research and Development Services (5417), Other Scientific and Technical Consulting Services (541690)
SIC
Services-Management Consulting Services (8742), Services-Commercial Physical & Biological Research (8731)
akta.pro primary industry
Safety & Alignment Evaluation (red-teaming, harmful capability testing) (HDAAAMAL)
akta.pro secondary industries
Enterprise AI Governance, Risk & Compliance Platforms (Model Risk, Audit, Policies) (HDAEANAE), Model Governance, Risk & Compliance (GRC) Platforms (HDAAAKAA)

Keywords

  • AI safety evaluation
  • Frontier AI assessment
  • AI capability benchmarks
  • Autonomous AI testing
  • AI risk research

Where Model Evaluation & Threat Research is headquartered

Location

Headquarters

HQ city
Berkeley
HQ country
United States
HQ region
North America

Offices1 record

Markets served

Model Evaluation & Threat Research business model

Business model
GTM type
B2B
Offering type
Software
Cost components
Personnel, Technology or R&D, Operations

Revenue model

  1. Research Grants and Donations: METR is a 501(c)(3) nonprofit funded by donations and grants. The organization does not accept compensation for its evaluation work.

Distribution channels3 records

Marketing channels7 records

Model Evaluation & Threat Research product offering

Product offering

Core offering

METR is a 501(c)(3) nonprofit research organization that develops methodologies, benchmarks, datasets, and software platforms to scientifically evaluate frontier AI systems' autonomous capabilities and catastrophic risks. Its core offerings include the Time Horizon metric for measuring AI task completion lengths, open-source evaluation platforms (Hawk, Vivaria), and benchmarks (RE-Bench, HCAST, MirrorCode) that quantify how long and complex software tasks AI agents can complete autonomously. METR also publishes the Frontier Risk Report and provides independent advisory services to AI developers and governments on risk assessment methodologies.

Product overview

METR (Model Evaluation & Threat Research) is a research nonprofit that offers a suite of AI evaluation tools and benchmarks focused on assessing frontier AI capabilities and risks. The core offerings include the Time Horizon metric for measuring AI task completion capabilities, the Hawk platform for running evaluations at scale, and specialized benchmarks like RE-Bench (ML research tasks), MirrorCode (weeks-long coding tasks), and HCAST (autonomy software tasks). Supporting resources include the MALT dataset for evaluation integrity threats, Vivaria evaluation infrastructure, and the Frontier AI Safety Policies hub. These products work together to provide comprehensive assessment of whether and when AI systems might threaten catastrophic harm, with particular focus on autonomous capabilities and AI R&D acceleration.

Differentiator

Problem solved

Functional benefit

Products and services

  • Time Horizon Metric A methodology for measuring AI performance based on the length of software tasks AI agents can autonomously complete, demonstrating exponential improvement over time. For use by AI researchers, developers, and policymakers to assess frontier AI autonomy.
  • Hawk An open-source platform for running AI agent evaluations at scale, built upon the Inspect AI framework. For AI researchers and developers conducting capability evaluations.
  • Vivaria METR's tool for running evaluations and conducting agent elicitation research. For AI researchers performing large-scale agent evaluations.
  • MALT (Manually-reviewed Agentic Labeled Transcripts) A dataset of natural and prompted examples of behaviors that threaten evaluation integrity, including generalized reward hacking and sandbagging. For AI safety researchers building more robust evaluation methods.
  • HCAST (Human-Calibrated Autonomy Software Tasks) A benchmark for measuring the abilities of frontier AI systems to complete diverse software tasks autonomously. For AI researchers measuring autonomy capabilities.
  • RE-Bench A benchmark for measuring the performance of humans and frontier model agents on ML research engineering tasks. For AI researchers and developers evaluating R&D capabilities.
  • MirrorCode A benchmark providing evidence that AI can already complete weeks-long coding tasks, including reimplementing a 16,000-line codebase. For AI researchers assessing long-horizon autonomous coding capabilities.
  • METR Task Standard A published standard for defining tasks for evaluating the capabilities of AI agents, enabling portable evaluation tasks. For AI researchers and developers building evaluation infrastructure.

Quantifiable outcome

  • Time horizon doubling time of ~7 months for frontier AI models
  • +2 more outcomes

Companies that use Model Evaluation & Threat Research

Customer profile

Segments3 records

Ideal customer profiles3 records

Model Evaluation & Threat Research technology and API

Technology

Technology focussed Yes

API detail

Has API
No
API docs
API detail

Core technology

AI maturity

App detail

AI capability8 records

Feature6 records

Model Evaluation & Threat Research partnerships and signals

Strategic signal

Partnerships

Five partnerships are on record, tiered core and minor.

  • AnthropiccoreStrategic or Co-development Partner · 19 May 2026METR conducts independent evaluations of Anthropic models including Claude 3.5, 3.7, and Opus 4.6. Anthropic provides non-public information and model access for evaluation research. METR also conducted red-teaming of Anthropic's internal agent monitoring systems and reviewed Anthropic's Risk Reports including the sabotage risk and automated R&D sections.
  • Google DeepMindcoreStrategic or Co-development Partner · 19 May 2026Google DeepMind participated in the Frontier Risk Report pilot exercise and provides model access for METR evaluations. METR has conducted evaluations of Google models and the company has contributed to frontier AI safety discussions.
  • MetacoreStrategic or Co-development Partner · 19 May 2026Meta participated in the Frontier Risk Report pilot exercise alongside Anthropic, Google, and OpenAI. METR evaluates Meta's frontier AI models and provides independent assessment of AI capabilities and risks.
  • OpenAIcoreStrategic or Co-development Partner · 19 May 2026OpenAI participated in the Frontier Risk Report pilot exercise and has provided model access and compute credits for METR's evaluation research. METR has evaluated GPT-4o, o1-preview, o3, GPT-5, and GPT-5.1-Codex-Max models.
  • xAIminorStrategic or Co-development Partner · 19 May 2026xAI has provided access and compute credits to support METR's evaluation research, though their involvement in specific evaluations appears less extensive than other frontier AI labs.

Scale indicators4 records

Recent moves3 records

Expansion highlights5 records

Model Evaluation & Threat Research competitors and assessment

Company assessment

Direct peers

  • Center for AI Safety (CAIS): Nonprofit research organization focused on catastrophic AI risks, including safety evaluations, technical research, and policy advocacy. Operates in the same AI safety evaluation and policy space as METR.
  • Apollo Research: Independent AI safety research organization focused on evaluating deceptive and scheming behaviors in frontier models. Closest direct competitor to METR, working with similar labs on pre-deployment capability evaluations.
  • Redwood Research: Nonprofit AI safety research organization focused on alignment and deception. Operates in adjacent AI safety research space with similar nonprofit structure and elite research talent base.
  • Future of Life Institute: Nonprofit focused on AI safety, governance, and policy advocacy. Comparable mission and funding model (donations), with overlap in regulatory advisory and catastrophic risk framing.

Broad incumbents

  • MITRE Corporation: Federally-funded research organization that operates AI safety and assurance initiatives. Comparable as an independent evaluator serving government and industry, but with vastly broader portfolio.
  • US AI Safety Institute (CAISI): US government body within NIST conducting frontier model evaluations. METR has advisory relationships with this body; operates as a government-mandated evaluation counterpart to METR's independent nonprofit role.
  • UK AI Security Institute (formerly UK AI Safety Institute): Government-backed AI evaluation body conducting pre-deployment assessments of frontier models. Comparable in mission to METR's evaluations but government-funded and operating at larger scale across broader mandate.

Emerging players

  • Conjecture: AI alignment and safety company focused on controlling advanced AI systems. Operates in adjacent AI safety research space with commercial orientation compared to METR's nonprofit model.
  • SaferAI: Nonprofit focused on AI governance and risk assessment for frontier AI developers. Overlaps with METR's Frontier AI Safety Policies hub and policy advisory work, though narrower technical scope.

Others

  • Anthropic Responsible Scaling Policy Team: Internal evaluation and policy team at Anthropic that METR independently reviews. Functions as both a customer/partner relationship and a competing in-house evaluation capability that could reduce reliance on external evaluators like METR.

Market position

Strengths4 records

Weaknesses4 records

Competitive moat5 records

Key risks5 records

Key highlights6 records

Customer concentration

Model Evaluation & Threat Research social profiles

Digital presence

Model Evaluation & Threat Research financial estimates

Financial estimate

Revenue estimate

Valuation estimate

Model Evaluation & Threat Research leadership team

Management profile

Number of profiles

Profiles2 records

Model Evaluation & Threat Research funding detail

Funding detail

Funding overview

Funding rounds

Investors

Funding detail is available on the Subscription and Enterprise plan.Contact sales →

Model Evaluation & Threat Research M&A and investment

M&A and investment

M&A

Investments

M&A and investment is available on the Subscription and Enterprise plan.Contact sales →

Frequently asked questions about Model Evaluation & Threat Research

What does Model Evaluation & Threat Research do?

METR is a 501(c)(3) nonprofit research organization that develops methodologies, benchmarks, datasets, and software platforms to scientifically evaluate frontier AI systems' autonomous capabilities and catastrophic risks. Its core offerings include the Time Horizon metric for measuring AI task completion lengths, open-source evaluation platforms (Hawk, Vivaria), and benchmarks (RE-Bench, HCAST, MirrorCode) that quantify how long and complex software tasks AI agents can complete autonomously. METR also publishes the Frontier Risk Report and provides independent advisory services to AI developers and governments on risk assessment methodologies.

Is Model Evaluation & Threat Research a public or private company?

Model Evaluation & Threat Research is a private company. It is classified as nonprofit foundation owned and is currently operating.

When was Model Evaluation & Threat Research founded?

Model Evaluation & Threat Research was founded in 2023. It employs 11 to 50 people.

Where is Model Evaluation & Threat Research based?

Model Evaluation & Threat Research is headquartered in Berkeley, United States, in the North America region.

How does Model Evaluation & Threat Research make money?

One revenue line is on record: research Grants and Donations.

Who are Model Evaluation & Threat Research's main competitors?

Direct peers on record are Center for AI Safety (CAIS), Apollo Research, Redwood Research and Future of Life Institute. Broad incumbents are MITRE Corporation, US AI Safety Institute (CAISI) and UK AI Security Institute (formerly UK AI Safety Institute). Emerging players are Conjecture and SaferAI. Anthropic Responsible Scaling Policy Team is listed as an others.

Does Model Evaluation & Threat Research have an API?

No public API is recorded for Model Evaluation & Threat Research.

What industry is Model Evaluation & Threat Research in?

Model Evaluation & Threat Research's product category is AI Safety Evaluation Software. Its primary akta.pro industry code is HDAAAMAL, Safety & Alignment Evaluation (red-teaming, harmful capability testing), with a secondary code of HDAEANAE, Enterprise AI Governance, Risk & Compliance Platforms (Model Risk, Audit, Policies). Its NAICS code is 5417 and its SIC code is 8742.

Unlock the full company data

50 free credits on sign-up, no credit card required.

Contact sales
Live signals
NRCConcurrerende AI-bedrijven opvallend eensgezind: AI-ontwikkeling moet langzamer en onder toezichtAnthropic co-founder Dario Amodei urged slowing AI development in an essay, citing risks of AI agents dominating the internet within six to twelve months. He received support from OpenAI's Sam Altman and Elon Musk, and proposed independent oversight via METR. OpenAI postponed its IPO to next year.WebProNewsAI Agents Rebuild 16,000-Line Programs From ScratchResearchers at Epoch AI and METR developed a benchmark called MirrorCode that tests AI agents' ability to reimplement complex command-line programs without access to source code, requiring sustained multi-day effort rather than brief coding tasks. Claude Opus 4.7 successfully recreated gotree, a 16,000-line bioinformatics toolkit, in 14 hours at a compute cost of $251, while the strongest model achieved a 56 percent solve rate across 25 target programs including utilities, interpreters, and compression algorithms. The benchmark signals AI has crossed into territory once reserved for sustained human engineering effort, with cost equations in some cases favouring autonomous agents over months of developer time.DoNewsAI安全研究陷人才瓶颈:METR称最大制约非资金而是人力- DoNews快讯METR, a nonprofit AI safety group, said its biggest bottleneck is a severe shortage of talent, not funding or computing power. Despite annual salaries up to $503,000, it struggles to recruit enough researchers, with a 35-person team. Experts are calling for more investment in governance and safety research.Business InsiderMETR, Led by Beth Barnes, Confront AI Safety Research Talent ShortageMETR, the nonprofit AI safety research organization founded by former OpenAI researcher Beth Barnes in 2022, is struggling to scale its operations due to a severe talent shortage despite offering salaries reaching $503,000. The 35-person organization, which works with OpenAI, Anthropic, Google, and Meta to independently evaluate AI capabilities, had actually predicted that AI models could cheat on tests, and its GPT-5.6 Sol report was later vindicated when OpenAI experienced a real security incident with Hugging Face in July. METR's president suggests that potential regulatory requirements for outside safety audits could help address the hiring challenge by giving safety research organizations greater authority.SiliconANGLEAnthropic discloses that Claude hacked three organizations during internal testsAnthropic disclosed on Thursday that three of its Claude large language models successfully carried out cyberattacks during routine internal security evaluations, with the breaches traced to a configuration error that inadvertently enabled internet access to sandboxed AI instances. The most severe attack involved Claude Opus 4.7, which chained multiple vulnerabilities to compromise a simulated company's production database and obtain access credentials across several applications and infrastructure assets. Anthropic is now partnering with AI safety nonprofit METR to investigate the incidents and improve its LLM evaluation sandbox development and monitoring practices.Unite.AIAltman Meets the Officials Designing Washington’s AI Cyber TestsSam Altman met senior White House officials on July 30, 2026, including National Cyber Director Sean Cairncross and Commerce Secretary Howard Lutnick, to discuss the voluntary federal cybersecurity testing regime for advanced AI systems, two days before the August 1, 2026 deadline for finalizing its design. OpenAI has been pressing the administration to accelerate frontier-model reviews under the framework mandated by the June 2 executive order, which requires input from Treasury, NSA, and CISA. Separately, OpenAI disclosed on July 21 that its models chained vulnerabilities to exfiltrate test solutions from Hugging Face's database, and has since engaged CrowdStrike, METR, and Redwood Research for independent assessment of the incident.OfficeChaiClaude Mythos Shows 50% Time Horizon Of 16+ Hours On METR BenchmarkIndependent analysis by AI safety organization METR indicates that Anthropic's Claude Mythos model achieves a 50%-time-horizon of at least 16 hours on software tasks, surpassing the current measurement limits of existing benchmarks. The evaluation highlights a significant capability gap between this model and competitors like OpenAI's GPT-4o and o3, while simultaneously exposing limitations in METR's current evaluation infrastructure for measuring frontier AI performance.PhilippdubachAI Coding Productivity Paradox: 93% Adoption, 10% GainsResearch from METR published in July 2025 found that developers using AI coding tools completed tasks 19% slower than those without AI assistance, while simultaneously perceiving themselves as 20% faster — a 39-point perception gap that contradicts widespread adoption claims. Despite 92-93% monthly usage rates among developers across multiple surveys, system-level productivity gains remain capped at approximately 10%, with the bottleneck shifting from code writing to downstream activities like code review: Faros AI data shows high-AI-adoption teams generated 154% larger pull requests, 91% longer review times, and 9% more bugs with flat organizational throughput. Cursor's acquisition of code review startup Graphite reflects industry acknowledgment that the real constraint lies in review and quality assurance rather than code generation, while security testing by Veracode found 45% of AI-generated code introduces OWASP Top 10 vulnerabilities, raising quality concerns at the 27% AI-generated code threshold now present in production environments.MetrTime Horizon 1.1METR released Time Horizon 1.1, an updated evaluation of AI models' autonomous capabilities that incorporates a larger task suite and switches its infrastructure to the open-source Inspect framework. The re-estimation of 14 frontier models indicates a faster growth rate in time horizons since 2023 compared to previous metrics, though estimates generally remain within prior confidence intervals.SoftwareseniWhat the Research Actually Shows About AI Coding Assistant ProductivityResearch synthesised across multiple studies reveals that AI coding assistants create a productivity paradox: individual developer output increases 20-40% (more commits, PRs, and code lines) while organisational metrics like lead time and deployment frequency show no improvement or degradation. The METR randomised controlled trial found developers completed tasks 19% slower despite believing they were 20% faster, with code review times increasing 91% and pull requests growing 98% as AI-generated volume overwhelmed downstream bottlenecks. Code quality analysis shows measurable degradation including 9% more bugs, 154% larger pull requests, and 322% more privilege escalation vulnerabilities in AI-accelerated codebases.