Written and maintained by CASRAI Editorial Board
Last updated
What METR is
METR — Model Evaluation & Threat Research, pronounced “meter” — is an independent research nonprofit that develops methods for assessing whether frontier AI systems pose catastrophic risks, and then applies those methods to evaluate specific models. It operates as a US 501(c)(3) organisation, funded by donations rather than by the AI companies whose systems it evaluates; METR states on its own site that it has not accepted funding from AI companies. Its stated mission is to “develop scientific methods to assess catastrophic risks stemming from AI systems’ autonomous capabilities and enable good decision-making about their development.”
METR is not a government body and has no statutory authority. It cannot block a model’s release. Its role is to produce independent, published findings that AI companies, policymakers and the public can use — most visibly through evaluation reports that some labs now attach to their own model announcements and system cards.
Where METR came from
METR began in 2022 as ARC Evals, the evaluations-focused offshoot of the Alignment Research Center (ARC), the nonprofit run by Paul Christiano. Beth Barnes, a former OpenAI alignment researcher, led that work. On 4 December 2023, ARC Evals spun out into its own independent nonprofit and took the name METR, with Barnes continuing as its lead. Christiano remained at ARC; he had initially been expected to join METR’s board but stepped back after taking a role at the US AI Safety Institute.
What METR evaluates
METR’s evaluations concentrate on autonomous capability — what an AI system can do without a human operator directing each step — rather than on general-purpose benchmark performance. Two lines of work anchor this:
- Autonomous replication and adaptation (ARA). METR uses this term for the capacity of an AI agent to acquire resources, create copies of itself, resist being shut down, and adapt to obstacles it encounters, without human assistance. ARA is the threat model behind what METR calls “rogue replication”: an agent establishing a persistent, self-sustaining presence outside its operator’s control. METR’s evaluations test for the component capabilities — making money, acquiring compute, installing and maintaining copies of itself, adapting to countermeasures — rather than for the end-to-end scenario itself.
- Time horizons. METR measures the length of software and research tasks a model can complete autonomously at a fixed success rate (typically 50%), expressed as a human-time equivalent. Its March 2025 report found this task-completion horizon growing roughly exponentially, with a doubling time on the order of seven months; an updated methodology, Time Horizon 1.1, followed in January 2026. This metric is now one of the standard reference points METR and others use to track how close a model is to the weeks-long sustained autonomy that an ARA scenario would require.
Alongside capability testing, METR studies behaviour that could undermine the evaluations themselves — models under-performing deliberately (“sandbagging”) or exploiting scoring loopholes (“reward hacking”) — and has published a dataset, MALT, documenting cases of this. It also runs RE-Bench, which compares model performance on machine-learning research-engineering tasks against human experts, as one input into its broader capability picture.
How the evaluations are run
METR’s evaluations are built around defined task suites with human performance baselines, so a model’s score can be compared to how long a skilled person takes to do the same work rather than judged in isolation. Because a model’s measured capability depends heavily on how well it is prompted and scaffolded, METR documents its elicitation methods explicitly and treats elicitation quality as a variable that has to be controlled for, not assumed away. It has also built and open-sourced tooling — including a platform called Hawk — for running these agent evaluations at scale across multiple models and task suites.
METR’s relationship to the labs it evaluates
METR is not employed by, or embedded inside, the AI companies it evaluates. It is a separate nonprofit that has arranged pre-deployment model access with several major developers — including OpenAI, Anthropic and Google DeepMind — specifically to run these evaluations before a model’s public release; its findings on models such as GPT-4o, o1-preview, o3, GPT-4.5, GPT-5 and Claude 3.5 and 3.7 Sonnet were produced under those arrangements. It has also evaluated models it did not have privileged early access to, such as DeepSeek’s V3 and R1, using public access after release.
This distinguishes METR from government evaluators working the same territory, such as the UK AI Security Institute or the US CAISI: METR is a privately funded nonprofit with no regulatory power, not a state body. It also distinguishes METR from a lab’s own internal safety team: METR is not party to the commercial incentives around a model’s launch date, and publishes its findings independently of whether a lab agrees with them.
METR and the frontier AI safety policy landscape
The idea of a published, capability-triggered safety policy — what Anthropic calls a Responsible Scaling Policy and other labs give other names — traces to METR: the concept was introduced by METR in 2023, and the first such policy was piloted that September. By METR’s own tracking, twelve AI companies have since published frontier AI safety policies of this kind, including Anthropic, OpenAI, Google DeepMind, Meta, Microsoft, Amazon, xAI, Cohere, NVIDIA, G42, Magic and Naver. METR’s evaluations are one of the mechanisms these policies rely on to decide whether a model has crossed a defined capability threshold.
Where to find METR’s published work
METR publishes its evaluation reports, methodology notes and research at metr.org/research, with detailed write-ups of individual model evaluations at evaluations.metr.org and shorter updates on its blog. Its findings on specific models are also frequently referenced directly in the releasing lab’s own system card for that model.
How this connects to NIKOLAI
CASRAI’s NIKOLAI dictionary of elements crosswalks the vocabulary used across frontier AI safety frameworks, regulatory text and evaluator standards — including METR’s — against a shared set of stable definitions. The concepts this page describes map most directly onto NIKOLAI’s N5 track (Evidence and Evaluations: evaluation, elicitation, saturation, evaluation-validity threats) and N8 track (Transparency and Review: external review, evaluator access, evaluator independence). As with the rest of NIKOLAI, these mappings are CASRAI’s own reading of METR’s published material; METR has not reviewed or endorsed them. Related CASRAI coverage of this evaluator ecosystem sits under the Third-Party Evaluation & Assurance subcluster, part of the wider Frontier AI Safety & Governance content cluster.
Frequently asked questions
Is METR a government agency or regulator?
No. METR is a privately funded 501(c)(3) nonprofit. It has no statutory authority and cannot require a lab to delay or withdraw a model; that is a meaningful difference from government evaluators such as the UK AI Security Institute or the US CAISI.
Who funds METR?
METR states that it is funded by donations and has not accepted funding from AI companies, which it treats as a condition of being able to evaluate those companies’ models independently.
What does “autonomous replication and adaptation” mean?
It is METR’s term for an AI agent’s ability to acquire resources, copy itself, resist shutdown and adapt to obstacles without human help — the set of capabilities that would let an agent sustain itself outside its operator’s control. METR evaluates for the component capabilities rather than waiting to observe the full scenario.
Does METR only evaluate OpenAI and Anthropic models?
No. METR has evaluated models from OpenAI, Anthropic, Google DeepMind and DeepSeek, among others, and its methodology is not tied to any one developer.
Has METR ever found a model too dangerous to release?
METR’s published evaluations to date — including of recent frontier models such as GPT-5 and GPT-5.1-Codex-Max — have found that self-improvement and “rogue replication” risks remain unlikely at current capability levels, while also noting that measured autonomous task-completion horizons are still increasing. METR does not have the authority to block a release either way; its findings feed into the releasing lab’s own decision under its safety policy.







