Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN3 · Thresholds and checkpointsProposednikolai-v0.1

Checkpoint Rule

NIKOLAI editorial proposal (unsourced, after Amodei): a checkpoint rule is a declared if-then rule of the form "if capability X, then certified properties Y and Z", naming the evidence types required, the certifier, and the consequence of failure to certify. This is Amodei's specific sense of "checkpoint" and is distinct from a training checkpoint (OpenAI PF) or a model-version checkpoint (GDM, xAI) — see the divergence note. This is element B8 of the source crosswalk.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

  • Amodei, "We Must Pace the Frontier" / Anthropic Risk Report, August 2026, PACE essay; RR Tables 3.10.A, 4.8.A

    "if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z"; evidence: "some combination of evaluations, interpretability analyses, and audits of training environments." Risk Report Tables 3.10.A and 4.8.A pair each threshold with "substantive standards."

    https://darioamodei.com/post/we-must-pace-the-frontier
  • Meta Advanced AI Scaling Framework v2, Table 1; §4.2.3

    Table 1: "Deploy with mitigations: Proceed with deployment of the Frontier AI only if sufficient mitigations are defined, implemented and validated to reduce risk to that of a moderate or lower model." Propensity acceptance criteria: "at least 40% on MASK and at most 50% on Agent Misalignment" (§4.2.3).

    https://ai.meta.com/static-resource/Meta_Advanced-AI-Scaling-Framework-v2

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. It is CASRAI's own reading of a published document. No lab, evaluator or regulator named here has declared, endorsed, or been consulted on this mapping. That will change only when an organisation files its own Mapping Declaration — see the non-endorsement policy.
OrganisationTheir term, as publishedMatchSource
Anthropic
Amodei essay / Risk Report / "The Urgency of Interpretability"
"if models have capability X, then they need to be accompanied by certifications of alignment properties Y and Z"; evidence: "some combination of evaluations, interpretability analyses, and audits of training environments" ("Pacing Within Democracies"). Risk Report Tables 3.10.A/4.8.A pair each threshold with "substantive standards"; CB-2's standard: "to the point where even well-resourced and -staffed threat actors would be unlikely to reliably jailbreak models or cause catastrophic harm via unauthorized access to or modification of models." "Interpretability analysis" is specified in "The Urgency of Interpretability" as a "'brain scan' checkup with a high chance of spotting deception, power-seeking, jailbreak flaws and cognitive strengths and weaknesses," used "in a loop with alignment training: diagnose, treat, scan again" — a proposed formalised interpretability test as the ASL-4 checkpoint's Y/Z-style certification.exact
confidence: high
Amodei, "We Must Pace the Frontier"
OpenAI
Preparedness Framework v2
PF Table 1 "risk-specific safeguard guidelines" per threshold, e.g. "Security controls meeting the High standard (App. C.3)"; for Critical, halt further development until Critical-standard safeguards are specified.
False friend embedded in the same source: PF's own word "checkpoint" means a training checkpoint, not this policy-rule sense — "we will select an appropriate checkpoint during development to be covered by the Preparedness Framework" (§3.2). Confirmed by the source document's own false-friends register ("checkpoint": Amodei's policy rule vs. OpenAI PF's training checkpoint vs. GDM/xAI's model-version checkpoint vs. Meta's two-stage evaluation stage). {PACE} {PF} {FSF} {FAIF26} {META}
close
confidence: high
OpenAI Preparedness Framework v2
Google DeepMind
Frontier Safety Framework v3.1
Risk acceptance conditions per T/CCL; for "ML R&D automation level 1": "We recommend Security level 4 for this capability threshold, but emphasize that this must be taken on by the frontier AI field as a whole" plus a safety case at any CCL (Table 3.2.2.a; glossary).
False friend embedded in the same source: GDM's own word "checkpoint" is used only for "various checkpoints or versions of a model" (s.1.3), not this policy-rule sense. Confirmed by the source document's own false-friends register.
close
confidence: high
Google DeepMind Frontier Safety Framework v3.1
xAI
xAI Frontier AI Framework (30 Jun 2026)
xAI's framework contains no capability-conditioned checkpoint rule. Its only use of "checkpoint" is "on various checkpoints or versions of a model" (s.2) — a model-version sense, unrelated to Amodei's policy-rule sense.
PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement disambiguating draft vs. final was found.
none
confidence: high
xAI Frontier AI Framework (30 Jun 2026, draft-marked)
Meta
Meta Advanced AI Scaling Framework v2
Table 1: "Deploy with mitigations: Proceed with deployment of the Frontier AI only if sufficient mitigations are defined, implemented and validated to reduce risk to that of a moderate or lower model." Propensity acceptance criteria attached to the Loss of Control "Capability checkpoint": "at least 40% on MASK and at most 50% on Agent Misalignment" (§4.2.3).
Meta's own "Capability checkpoint" names the first stage of a two-stage Loss of Control evaluation, a further sense distinct from Amodei's policy-rule sense, OpenAI's training-checkpoint sense, and GDM/xAI's model-version sense — confirmed by the source document's own false-friends register.
close
confidence: high
Meta Advanced AI Scaling Framework v2
METR
METR Common Elements
Capability Thresholds "require new robust mitigations" (/common-elements).close
confidence: medium
METR (common-elements)
Frontier Model Forum
FMF Risk Taxonomy and Thresholds
"If-then" structure (above); "Acceptable Development or Deployment Thresholds": "Criteria that determine whether a model that has crossed an enabling capability threshold can be safely deployed or trained further after implementing safeguards." (s3.1)close
confidence: high
FMF Risk Taxonomy and Thresholds
Safety Framework Cards (discovery)
Safety Framework Cards (SSRN 7061798)
"governance triggers" dimension named in discovery sweep; full text paywalled/unread.
Unverified: full text inaccessible to the source document's own research pass.
none
confidence: low
Discovery sweep — Safety Framework Cards (SSRN 7061798, paywalled/unread)
Magic
Magic AGI Readiness Policy
"If we have not developed adequate dangerous capability evaluations by the time these benchmark thresholds are exceeded, we will halt further model development until our dangerous capability evaluations are ready."close
confidence: high
Magic AGI Readiness Policy
G42
G42 Frontier Safety Framework
G42 pairs each threshold with a required Deployment Mitigation Level and Security Mitigation Level that "must be achieved before the capability threshold is reached."close
confidence: medium
G42 Frontier Safety Framework

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • California SB 53

    Frameworks must describe "(3) Applying mitigations to address the potential for catastrophic risks based on the results of assessments" (22757.12(a)) — a general obligation to have mitigations, with no rule content of the if-then, certified-properties kind. Scored RL in the source document.

    California SB 53
  • US EO 14409 / UK Seoul Commitments (GOV)

    EO 14409 links designation to a voluntary access framework (Sec. 3(b)); Seoul Commitment IV: developers "develop and deploy their systems and models only if they assess that residual risks would stay below the thresholds" — general commitments, not a stated if-then checkpoint rule. Scored RL in the source document.

    Executive Order 14409 / Seoul Frontier AI Safety Commitments

Divergence

Where sources materially disagree

The label "checkpoint" is a false-friend collision confirmed by the source document's own false-friends register: Amodei's essay uses it for a capability-conditioned policy rule (if capability X, then certified properties Y and Z) — the concept this NIKOLAI element models; OpenAI's Preparedness Framework uses it for a training checkpoint selected for Framework coverage; GDM and xAI both use it only for "checkpoints or versions of a model" in the release lifecycle, and xAI's framework contains no capability-conditioned rule at all; and Meta's "Capability checkpoint" names the first stage of its own two-stage Loss of Control evaluation. NIKOLAI's checkpoint-rule element models Amodei's sense only. {PACE} {PF} {FSF} {FAIF26} {META}

Gap

No source names a certifier other than the developer itself, and only Amodei requires alignment properties (not just safeguards) as the consequent. Meta's propensity acceptance criteria are the nearest operational example.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →