Source of record
Where this definition comes from
Amodei, "We Must Pace the Frontier" / Anthropic Risk Report, August 2026, PACE essay; RR Tables 3.10.A, 4.8.A
“"if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z"; evidence: "some combination of evaluations, interpretability analyses, and audits of training environments." Risk Report Tables 3.10.A and 4.8.A pair each threshold with "substantive standards."”
https://darioamodei.com/post/we-must-pace-the-frontierMeta Advanced AI Scaling Framework v2, Table 1; §4.2.3
“Table 1: "Deploy with mitigations: Proceed with deployment of the Frontier AI only if sufficient mitigations are defined, implemented and validated to reduce risk to that of a moderate or lower model." Propensity acceptance criteria: "at least 40% on MASK and at most 50% on Agent Misalignment" (§4.2.3).”
https://ai.meta.com/static-resource/Meta_Advanced-AI-Scaling-Framework-v2
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Amodei essay / Risk Report / "The Urgency of Interpretability" | “"if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z"; evidence: "some combination of evaluations, interpretability analyses, and audits of training environments" ("Pacing Within Democracies"). Risk Report Tables 3.10.A/4.8.A pair each threshold with "substantive standards"; CB-2's standard: "to the point where even well-resourced and -staffed threat actors would be unlikely to reliably jailbreak models or cause catastrophic harm via unauthorized access to or modification of models." "Interpretability analysis" is specified in "The Urgency of Interpretability" as a "'brain scan' checkup with a high chance of spotting deception, power-seeking, jailbreak flaws and cognitive strengths and weaknesses," used "in a loop with alignment training: diagnose, treat, scan again" — a proposed formalised interpretability test as the ASL-4 checkpoint's Y/Z-style certification.” | exact confidence: high | Amodei, "We Must Pace the Frontier" |
| OpenAI Preparedness Framework v2 | “PF Table 1 "risk-specific safeguard guidelines" per threshold, e.g. "Security controls meeting the High standard (App. C.3)"; for Critical, halt further development until Critical-standard safeguards are specified.” False friend embedded in the same source: PF's own word "checkpoint" means a training checkpoint, not this policy-rule sense — "we will select an appropriate checkpoint during development to be covered by the Preparedness Framework" (§3.2). Confirmed by the source document's own false-friends register ("checkpoint": Amodei's policy rule vs. OpenAI PF's training checkpoint vs. GDM/xAI's model-version checkpoint vs. Meta's two-stage evaluation stage). {PACE} {PF} {FSF} {FAIF26} {META} | close confidence: high | OpenAI Preparedness Framework v2 |
| Google DeepMind Frontier Safety Framework v3.1 | “Risk acceptance conditions per T/CCL; for "ML R&D automation level 1": "We recommend Security level 4 for this capability threshold, but emphasize that this must be taken on by the frontier AI field as a whole" plus a safety case at any CCL (Table 3.2.2.a; glossary).” False friend embedded in the same source: GDM's own word "checkpoint" is used only for "various checkpoints or versions of a model" (s.1.3), not this policy-rule sense. Confirmed by the source document's own false-friends register. | close confidence: high | Google DeepMind Frontier Safety Framework v3.1 |
| xAI xAI Frontier AI Framework (30 Jun 2026) | “xAI's framework contains no capability-conditioned checkpoint rule. Its only use of "checkpoint" is "on various checkpoints or versions of a model" (s.2) — a model-version sense, unrelated to Amodei's policy-rule sense.” PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement disambiguating draft vs. final was found. | none confidence: high | xAI Frontier AI Framework (30 Jun 2026, draft-marked) |
| Meta Meta Advanced AI Scaling Framework v2 | “Table 1: "Deploy with mitigations: Proceed with deployment of the Frontier AI only if sufficient mitigations are defined, implemented and validated to reduce risk to that of a moderate or lower model." Propensity acceptance criteria attached to the Loss of Control "Capability checkpoint": "at least 40% on MASK and at most 50% on Agent Misalignment" (§4.2.3).” Meta's own "Capability checkpoint" names the first stage of a two-stage Loss of Control evaluation, a further sense distinct from Amodei's policy-rule sense, OpenAI's training-checkpoint sense, and GDM/xAI's model-version sense — confirmed by the source document's own false-friends register. | close confidence: high | Meta Advanced AI Scaling Framework v2 |
| METR METR Common Elements | “Capability Thresholds "require new robust mitigations" (/common-elements).” | close confidence: medium | METR (common-elements) |
| Frontier Model Forum FMF Risk Taxonomy and Thresholds | “"If-then" structure (above); "Acceptable Development or Deployment Thresholds": "Criteria that determine whether a model that has crossed an enabling capability threshold can be safely deployed or trained further after implementing safeguards." (s3.1)” | close confidence: high | FMF Risk Taxonomy and Thresholds |
| Safety Framework Cards (discovery) Safety Framework Cards (SSRN 7061798) | “"governance triggers" dimension named in discovery sweep; full text paywalled/unread.” Unverified: full text inaccessible to the source document's own research pass. | none confidence: low | Discovery sweep — Safety Framework Cards (SSRN 7061798, paywalled/unread) |
| Magic Magic AGI Readiness Policy | “"If we have not developed adequate dangerous capability evaluations by the time these benchmark thresholds are exceeded, we will halt further model development until our dangerous capability evaluations are ready."” | close confidence: high | Magic AGI Readiness Policy |
| G42 G42 Frontier Safety Framework | “G42 pairs each threshold with a required Deployment Mitigation Level and Security Mitigation Level that "must be achieved before the capability threshold is reached."” | close confidence: medium | G42 Frontier Safety Framework |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- California SB 53
Frameworks must describe "(3) Applying mitigations to address the potential for catastrophic risks based on the results of assessments" (22757.12(a)) — a general obligation to have mitigations, with no rule content of the if-then, certified-properties kind. Scored RL in the source document.
California SB 53 - US EO 14409 / UK Seoul Commitments (GOV)
EO 14409 links designation to a voluntary access framework (Sec. 3(b)); Seoul Commitment IV: developers "develop and deploy their systems and models only if they assess that residual risks would stay below the thresholds" — general commitments, not a stated if-then checkpoint rule. Scored RL in the source document.
Executive Order 14409 / Seoul Frontier AI Safety Commitments
Divergence
Where sources materially disagree
The label "checkpoint" is a false-friend collision confirmed by the source document's own false-friends register: Amodei's essay uses it for a capability-conditioned policy rule (if capability X, then certified properties Y and Z) — the concept this NIKOLAI element models; OpenAI's Preparedness Framework uses it for a training checkpoint selected for Framework coverage; GDM and xAI both use it only for "checkpoints or versions of a model" in the release lifecycle, and xAI's framework contains no capability-conditioned rule at all; and Meta's "Capability checkpoint" names the first stage of its own two-stage Loss of Control evaluation. NIKOLAI's checkpoint-rule element models Amodei's sense only. {PACE} {PF} {FSF} {FAIF26} {META}
Gap
No source names a certifier other than the developer itself, and only Amodei requires alignment properties (not just safeguards) as the consequent. Meta's propensity acceptance criteria are the nearest operational example.







