Written and maintained by CASRAI Editorial Board
Last updated
Last verified: September 20, 2026. “Capability threshold” is one of the most load-bearing phrases in frontier AI governance — it is the trigger that turns a lab’s safety framework from a document into an obligation. Almost every major lab, regulator, and evaluator now uses some version of the term. Almost none of them mean the same thing by it, and most don’t say what the number actually is. This guide walks through fourteen organizations’ published language on capability thresholds, sourced from CASRAI’s own NIKOLAI crosswalk on the term, and verified independently against primary documents where those documents are public.
The short version, stated plainly: SB 53 mandates thresholds but provides no definition. xAI cites thresholds without publication. Amazon uses undefined qualitative terms like “material uplift” and “reliably.” Magic alone discloses a numeric public threshold — a single benchmark score, with a named comparison point. Everyone else sits somewhere between those two poles.
What a “Capability Threshold” Is Supposed to Do
Across every framework in this piece, a capability threshold plays the same structural role: it is the line a model’s tested abilities have to cross before a specific obligation kicks in — new safeguards, a halt to further development, a mandatory disclosure, or a deployment restriction. The idea is simple. The execution is not. A threshold can be stated as a precise benchmark score, a qualitative description of what a model can help someone do, a requirement to define one without saying what it is, or a classified number nobody outside government sees. All four of those are currently in use, sometimes by the same organization for different risk domains.
The Labs: Thresholds That Exist, Stated With Varying Precision
Anthropic — Two Named Thresholds, Neither One a Number
Anthropic’s Responsible Scaling Policy defines its “AI R&D threshold (current)” with unusual precision for a qualitative test. Per Anthropic’s Risk Report, August 2026 (RSP v3.4, §1.3.1, §3.2), the threshold is met if “(1) our models would be able to fully substitute for our entire set of Research Scientists and Research Engineers, at competitive costs (i.e., within a factor of 5); or (2) there is ‘dramatic acceleration’ of the pace of AI progress for reasons that likely relate to the automation of AI R&D.” That’s a real operational test — a 5x cost ceiling, a defined comparison to a documented baseline rate of progress — even though it isn’t a single benchmark number. The same Risk Report defines a separate, distinct CB-2 threshold for chemical and biological weapons: AI systems that “can functionally substitute for the scarce human expertise that is currently the primary barrier to novel development of chemical and biological weapons with potential for catastrophic harm” (§1.3.2). Anthropic has revised the AI R&D threshold’s wording twice since its prior Risk Report (in RSP v3.1 and v3.4), each time tightening the operational criteria — itself a small case study in how much a “defined” threshold can still move.
OpenAI — Thresholds Used, Never Formally Defined at the Term Level
OpenAI’s Preparedness Framework v2 (§2.2) states: “Capability thresholds concretely describe things an AI system might be able to help someone do or might be able to do on its own that could meaningfully increase risk of severe harm.” The Framework then does define concrete High and Critical thresholds per Tracked Category — for AI Self-improvement, for instance, the High threshold is a model whose impact is “equivalent to giving every OpenAI researcher a highly performant mid-career research engineer assistant,” and the Critical threshold is a model “capable of recursively self improving,” defined either as a superhuman research-scientist agent or as compressing a generational model improvement into one-fifth the wall-clock time, sustained for months. Those are substantive, if qualitative, tests — but the term “capability threshold” itself is used throughout the document without a standalone formal definition separate from the category-specific tables.
Google DeepMind — Two Tiers, Both Formally Defined
DeepMind’s Frontier Safety Framework v3.1 (glossary, p.18) is the most terminologically precise of the three major labs. It defines Critical Capability Levels (CCLs) as “the main capability thresholds around which we have built the Framework process” — the levels at which, absent mitigation, a model “may pose heightened risk of severe harm.” It defines Tracked Capability Levels (TCLs) as “capability thresholds which capture a lower level of risks than our CCLs” — heightened risk of “significant but not severe” harm. Concrete examples exist for CBRN, cyber, harmful manipulation, and ML R&D domains, each with a named security-level recommendation attached. FSF v3.1 also introduces “alert thresholds” — set marginally earlier than a CCL, to flag when one may be reached before the next assessment cycle. Of the three major labs, DeepMind is the only one that formally defines the term “capability threshold” itself, inside a named glossary, rather than leaving it to be inferred from category tables.
xAI — The Term Without the Number
xAI’s draft Frontier AI Framework (30 June 2026) and its Grok 4.6 model card both cite “FAIF safety thresholds” as the operative mechanism — but neither document states what those thresholds actually are. The June 2026 framework discusses “analyzing whether a threshold has been reached” (§2.3) without stating one. The PDF’s own metadata still labels it a “Privileged/Confidential DRAFT working FRAMEWORK DOC,” and no public xAI statement has since confirmed the June 2026 version is final rather than draft. xAI’s earlier frameworks (December 2025 FAIF, August 2025 RMF) used a different mechanism entirely — deployment acceptance criteria stated as pass/fail rates, such as “maintaining an answer rate of less than 1 out of 20 on restricted queries” — rather than capability thresholds proper. xAI is the case where the vocabulary of thresholds is fully present and the actual values are not.
The Regulators: One Measurable Requirement, Two That Mandate Without Defining, One Classified
The EU — Measurable, but Not Shared Across Signatories
The EU GPAI Code of Practice’s Safety and Security chapter, Measure 4.1, is the most procedurally rigorous requirement in this whole set. It requires each Signatory to “define appropriate systemic risk tiers that: (i) are defined in terms of model capabilities, and may additionally incorporate model propensities, risk estimates, and/or other suitable metrics; (ii) are measurable; and (iii) comprise at least one systemic risk tier that has not been reached by the model.” That’s a real, checkable standard — measurability is mandatory, not aspirational. What Measure 4.1 does not do is specify a shared scale: each Signatory defines its own tiers, so “systemic risk tier 2” at one lab and at another are not directly comparable numbers, even though both are individually measurable. See GPAI Systemic Risk: The EU AI Act Term Explained for the fuller mechanics of how those tiers interact with the Act’s broader risk classification.
California SB 53 — Required by Statute, Defined by No One
California’s SB 53, §22757.12(a)(2), requires large frontier developers’ frameworks to describe “defining and assessing thresholds used by the large frontier developer to identify and assess whether a frontier model has capabilities that could pose a catastrophic risk, which may include multiple-tiered thresholds.” The statute mandates that a developer have thresholds and disclose how it defines them. It does not itself define the word “threshold” anywhere in the bill text — the definitional work is delegated entirely to each developer’s own framework. This is the cleanest example in the corpus of a legal requirement to threshold without a corresponding legal definition of what a threshold is.
US Executive Order 14409 and the Seoul Commitments — Classified, and Outcome-Gated Rather Than Capability-Based
Executive Order 14409 designates a cyber-domain threshold whose value is classified — a case where a threshold plainly exists and is used by government to make decisions, but is by design unavailable for public comparison against anyone else’s. The Seoul Frontier AI Safety Commitments take a different tack from either publication or classification: Commitment II defines thresholds as the point “at which severe risks posed by a model or system, unless adequately mitigated, would be deemed intolerable.” That framing is narrower than a capability threshold proper — it gates on the outcome (intolerable risk after mitigation) rather than on a specific tested capability, which makes it structurally closer to a risk-acceptance criterion than to the capability-first thresholds labs publish. See Seoul Frontier AI Safety Commitments for the full signatory list and pledge text.
The Evaluators and Coalitions: Naming the Concept, Not Setting the Number
METR’s Common Elements document defines “Capability Thresholds” generically as “thresholds at which specific AI capabilities would pose severe risk and require new mitigations” — establishing the term as a recommended element of a good framework, without prescribing values, since METR’s role is evaluating labs’ frameworks rather than publishing its own capability limits. The Frontier Model Forum’s Risk Taxonomy and Thresholds report (§3.1) defines “Enabling Capability Thresholds” — also called “critical capability levels” or “capability thresholds” in the same document — as “abilities that could potentially enable extreme harms if the model is deployed without additional safeguards,” illustrated with an example like PhD-level biology proficiency in a threat-relevant context. Both organizations are doing taxonomy work: naming and standardizing the concept industry-wide, which is a different job from any single lab publishing its own trigger values.
Amazon: Qualitative and Uplift-Based, by Design
Amazon’s Frontier Model Safety Framework defines “Critical Capability Thresholds” as “a set of model capabilities that have the potential to cause significant harm to the public if misused,” with one threshold per Critical Risk Domain — CBRN, Offensive Cyber Operations, Harmful Manipulation, and Loss of Control. Each is written as an uplift-based description rather than a benchmark score: the CBRN threshold, for instance, is met when a model is “capable of providing expert-level, interactive instruction that provides material uplift (beyond other publicly available models in known harnesses) that would enable a non-subject matter expert to reliably produce and deploy a CBRN weapon.” “Material uplift” and “reliably” are doing real work in that sentence, and neither term is given a numeric or independently-verifiable operational definition elsewhere in the document — a structural choice common to uplift-study methodology generally, not unique to Amazon, but one that leaves the actual trigger point to internal judgment call rather than external replication.
Magic: The One Disclosed Number in the Entire Corpus
Magic’s AGI Readiness Policy stands alone. It states plainly: the commitment triggers “when, at the end of a training run, our models exceed a threshold of 50% accuracy on LiveCodeBench” (Pass@1) — and names a contemporaneous public-model comparison point directly in the policy, Claude-3.5-Sonnet’s documented 48.8% on the same benchmark at the time. That is a complete, independently checkable threshold: a named public benchmark, a specific percentage, and a stated baseline for context. Nothing else in this corpus does all three. It is a narrower commitment than Anthropic’s or DeepMind’s multi-domain frameworks — LiveCodeBench is one coding benchmark, not a CBRN or cyber-uplift test — but within its scope, it is the only threshold here a third party could reproduce from the public policy document alone.
Related but Not Mapped: Meta’s Outcome-Based “Risk Thresholds”
Meta’s Advanced AI Scaling Framework v2 (Appendix I) defines “Risk thresholds” as “the incremental levels of risk that a Frontier AI model might pose towards realization of a catastrophic outcome.” NIKOLAI scores this as related to, but not a direct mapping of, the capability-threshold element — Meta’s “Risk thresholds” are outcome-based risk levels, and capability enters Meta’s framework separately, through its own distinctly-named “Enabling capabilities” and cyber-domain “Operational threshold” concepts. It is included here for completeness, and as a reminder that similar-sounding terms across frameworks do not always describe the same thing — part of why NIKOLAI tracks match type and confidence separately for every row rather than treating a shared word as a shared definition.
What the Fourteen Rows Add Up To
Put the crosswalk side by side and a pattern falls out immediately: SB 53 mandates thresholds but provides no definition. xAI cites thresholds without publication. Amazon uses undefined qualitative terms — “material uplift,” “reliably” — that are never numerically operationalized. Magic alone discloses numeric public thresholds. Everyone else in this piece sits somewhere on the spectrum between those poles: Anthropic and OpenAI publish real operational tests that are not single benchmark numbers; DeepMind formally defines the term itself and attaches named security levels to each tier; the EU mandates measurability without a shared scale; METR and the Frontier Model Forum standardize the vocabulary without setting values; EO 14409’s number exists but is classified; and the Seoul Commitments gate on outcome rather than capability at all.
None of this means the qualitative frameworks are worse-designed than Magic’s. A single numeric threshold is easiest to verify from the outside, but it is also the easiest to game or to render meaningless as a proxy once a model is optimized against exactly that benchmark — which is part of why most multi-domain frameworks lean on uplift studies, expert red-teaming, and holistic judgment instead of a single score. The point of laying the fourteen rows out together isn’t to declare a winner; it’s that the same two words — “capability threshold” — currently point at four structurally different kinds of commitment, and a reader comparing two frameworks by title alone has no way to know which kind they’re getting.
Where NIKOLAI Fits In
This page is a direct deep-dive on NIKOLAI’s capability-threshold element (element B4 in the source crosswalk), part of Track N3, Thresholds and Checkpoints, in CASRAI’s frontier-AI-safety dictionary. Every row above — Anthropic through Meta — reflects only what that organization has explicitly published; NIKOLAI calls this a “shadow mapping” unless the organization itself has gone through NIKOLAI’s Mapping Declarations process to confirm how the term maps to its own usage. As of this writing, none of the fourteen organizations covered here has made that declaration, so every mapping above is CASRAI’s own read of the public record, not a confirmation from Anthropic, OpenAI, Google DeepMind, xAI, the EU, California, the U.S. government, METR, the Frontier Model Forum, Amazon, Magic, or Meta. NIKOLAI is CASRAI’s own independent reference work — unendorsed by any of them.
NIKOLAI’s own gap note on this element proposes a structural fix for exactly the inconsistency this page documents: a threshold.disclosureStatus property, valued as quantified, qualitative, referenced-but-undefined, or classified. Tag every organization’s threshold language against those four values, and the crosswalk above sorts itself cleanly: Magic alone lands on quantified; Anthropic, OpenAI, DeepMind, the EU, METR, FMF, and Amazon land on qualitative (each with very different levels of operational precision within that bucket); xAI and SB 53 land on referenced-but-undefined; EO 14409 lands on classified. That single field would not resolve the underlying disagreement about what a threshold should be — but it would make the disagreement itself visible and comparable, which is the more modest goal NIKOLAI is built around.
A Note on Scope: This Is Not the RSP Comparison
This page is deliberately narrower than, and non-duplicative of, Responsible Scaling Policy: What It Is and How the Major Labs Compare, which covers Anthropic’s, OpenAI’s, and Google DeepMind’s scaling-policy mechanics as a whole — governance structure, review cadence, escalation process — across those three labs only. That guide does not cover xAI, the EU, SB 53, METR, the Frontier Model Forum, Amazon, or Magic at all, and it does not isolate the capability-threshold concept specifically from the rest of each policy. This page does the opposite: it isolates one term, across all fourteen organizations that use it, labs and regulators and evaluators alike, regardless of how the rest of their frameworks are structured.
Frequently Asked Questions
Which organization has the most precisely defined capability threshold?
Magic is the only organization in this corpus with a fully disclosed, numeric, independently checkable public threshold — 50% accuracy on LiveCodeBench (Pass@1), with a named comparison baseline. Google DeepMind is the only major lab to formally define the term “capability threshold” itself (as CCLs and TCLs) inside a named glossary, though its actual trigger values are qualitative descriptions rather than single benchmark scores.
Does California SB 53 define what a capability threshold is?
No. SB 53 §22757.12(a)(2) requires large frontier developers to describe how they define and assess thresholds for catastrophic risk, but the statute itself does not define the term “threshold” — that work is left entirely to each developer’s own framework.
Has xAI published its capability thresholds?
Not as of its 30 June 2026 draft Frontier AI Framework or its Grok 4.6 model card, both of which cite “FAIF safety thresholds” without stating values. The June 2026 framework’s own PDF metadata still labels it a confidential draft. xAI’s earlier frameworks used pass/fail deployment acceptance criteria rather than capability thresholds as such.
What is NIKOLAI’s proposed fix for the inconsistent use of “capability threshold”?
NIKOLAI proposes adding a threshold.disclosureStatus property to any threshold record, valued as quantified, qualitative, referenced-but-undefined, or classified — making each organization’s actual level of disclosure explicit and comparable, rather than treating the shared phrase “capability threshold” as evidence of a shared standard.
Is this the same as the Responsible Scaling Policy comparison guide?
No. See Responsible Scaling Policy: What It Is and How the Major Labs Compare for RSP mechanics across Anthropic, OpenAI, and Google DeepMind specifically. This page instead isolates the capability-threshold term across all fourteen organizations in NIKOLAI’s crosswalk, including regulators, evaluators, and labs that guide does not cover.







