Source of record
Where this definition comes from
Anthropic Risk Report, August 2026, §1.3.1, §1.3.2, §3.2
“"AI R&D threshold (current)": met if "(1) our models would be able to fully substitute for our entire set of Research Scientists and Research Engineers, at competitive costs (i.e., within a factor of 5); or (2) there is 'dramatic acceleration' of the pace of AI progress for reasons that likely relate to the automation of AI R&D" (§1.3.1, §3.2).”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfGoogle DeepMind Frontier Safety Framework v3.1, glossary, p.18
“"Critical Capability Levels (CCLs): are the main capability thresholds around which we have built the Framework process." "Tracked Capability Levels (TCLs): are capability thresholds which capture a lower level of risks than our CCLs." (glossary, p.18)”
https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3-1.pdfMagic AGI Readiness Policy, Threshold Definition
“The sole disclosed numeric public capability threshold in the corpus — "when, at the end of a training run, our models exceed a threshold of 50% accuracy on LiveCodeBench" (Pass@1), with the contemporaneous public-model baseline "Claude-3.5-Sonnet 48.8%" given for comparison.”
https://magic.dev/agi-readiness-policy
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report, August 2026 / RSP v3.4 | “"AI R&D threshold (current)": met if "(1) our models would be able to fully substitute for our entire set of Research Scientists and Research Engineers, at competitive costs (i.e., within a factor of 5); or (2) there is 'dramatic acceleration' of the pace of AI progress for reasons that likely relate to the automation of AI R&D" (§1.3.1, §3.2). CB-2 threshold: "AI systems that can functionally substitute for the scarce human expertise that is currently the primary barrier to novel development of chemical and biological weapons with potential for catastrophic harm ..." (§1.3.2). RSP §1 table header: "Capability or usage threshold."” Source document's own citation for the RSP §1 table-header note also carried a second bracket tag ({DAN}) that does not resolve against NIKOLAI's provided sources table; omitted here rather than guessed at. | exact confidence: high | Anthropic Risk Report, August 2026 |
| OpenAI Preparedness Framework v2 / Frontier Governance Framework | “"Capability thresholds: concretely describe things an AI system might be able to help someone do or might be able to do on its own that could meaningfully increase risk of severe harm." (§2.2). FGF: capability threshold is "used but never formally defined"; "Risk tiers" are the published scale (§2.2, §2.3).” | exact confidence: high | OpenAI Preparedness Framework v2 |
| Google DeepMind Frontier Safety Framework v3.1 | “"Critical Capability Levels (CCLs): are the main capability thresholds around which we have built the Framework process." "Tracked Capability Levels (TCLs): are capability thresholds which capture a lower level of risks than our CCLs." (glossary, p.18)” | exact confidence: high | Google DeepMind Frontier Safety Framework v3.1 |
| xAI xAI Frontier AI Framework (30 Jun 2026) / Grok 4.6 card / earlier xAI frameworks | “June 2026 framework: "When analyzing whether a threshold has been reached ..." (s.2.3), with no thresholds stated. Grok 4.6 card: "FAIF safety thresholds" cited with no values given. The earlier Dec 2025 FAIF and Aug 2025 RMF used deployment acceptance criteria instead: "maintaining an answer rate of less than 1 out of 20 on restricted queries" and "a dishonesty rate of less than 1 out of 2 on MASK."” PDF metadata /Title of the 30 Jun 2026 framework reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement disambiguating draft vs. final was found. | none confidence: medium | xAI Frontier AI Framework (30 Jun 2026, draft-marked) |
| EU EU GPAI Code of Practice, Safety and Security chapter | “Measure 4.1: for each identified systemic risk, Signatories "define appropriate systemic risk tiers that: (i) are defined in terms of model capabilities, and may additionally incorporate model propensities, risk estimates, and/or other suitable metrics; (ii) are measurable; and (iii) comprise at least one systemic risk tier that has not been reached by the model." Content of tiers left to each Signatory; no shared tier scale published in the chapter.” | exact confidence: high | EU GPAI Code of Practice, Safety and Security chapter |
| California SB 53 California SB 53 | “Frameworks must describe "(2) Defining and assessing thresholds used by the large frontier developer to identify and assess whether a frontier model has capabilities that could pose a catastrophic risk, which may include multiple-tiered thresholds." (22757.12(a)). "Threshold" itself is not defined by the statute.” | none confidence: medium | California SB 53 |
| US Government (Executive Order 14409) EO 14409 / Seoul Frontier AI Safety Commitments | “EO 14409 designation threshold is classified (cyber domain). Seoul Commitment II: thresholds are "at which severe risks posed by a model or system, unless adequately mitigated, would be deemed intolerable" — a narrower, outcome-gated framing than a capability threshold proper.” | narrow confidence: medium | Executive Order 14409 |
| METR METR Common Elements | “"Capability Thresholds": "Thresholds at which specific AI capabilities would pose severe risk and require new mitigations."” | exact confidence: high | METR (common-elements) |
| Frontier Model Forum FMF Risk Taxonomy and Thresholds | “"Enabling Capability Thresholds (also called 'critical capability levels' or 'capability thresholds'): Abilities that could potentially enable extreme harms if the model is deployed without additional safeguards." (s3.1, p.10)” | exact confidence: high | FMF Risk Taxonomy and Thresholds |
| Safety Framework Cards (discovery) Safety Framework Cards (SSRN 7061798) | “"capability thresholds" dimension named in discovery sweep; full text paywalled/unread.” Unverified: full text inaccessible to the source document's own research pass. | none confidence: low | Discovery sweep — Safety Framework Cards (SSRN 7061798, paywalled/unread) |
| Amazon Amazon Frontier Model Safety Framework | “"Critical Capability Thresholds": "a set of model capabilities that have the potential to cause significant harm to the public if misused" (Overview, p.1); one qualitative threshold per domain (CBRN, Offensive Cyber Operations, Automated AI R&D), stated as an uplift-based description rather than a benchmark score.” | exact confidence: high | Amazon Frontier Model Safety Framework |
| Magic Magic AGI Readiness Policy | “The sole disclosed numeric public capability threshold in the corpus — "when, at the end of a training run, our models exceed a threshold of 50% accuracy on LiveCodeBench" (Pass@1), with the contemporaneous public-model baseline "Claude-3.5-Sonnet 48.8%" given for comparison (Threshold Definition).” | exact confidence: high | Magic AGI Readiness Policy |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- Meta
"Risk thresholds: are the incremental levels of risk that a Frontier AI model might pose towards realization of a catastrophic outcome" (Appendix I) — these are outcome-based risk levels, not capability thresholds; capability enters only through separately-named "Enabling capabilities" and "Operational threshold" (cyber). Scored RL (related, not a mapping) in the source document.
Meta Advanced AI Scaling Framework v2
Gap
SB 53 requires thresholds but defines none; xAI's current framework and the Grok 4.6 card cite thresholds that are not published; Amazon's thresholds are qualitative only, with undefined terms such as "material uplift" and "reliably." Magic is the one exception with a disclosed numeric public threshold. NIKOLAI needs a `threshold.disclosureStatus` value set (quantified; qualitative; referenced-but-undefined; classified).







