Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN3 · Thresholds and checkpointsProposednikolai-v0.1

Threshold Status

NIKOLAI editorial proposal (unsourced): threshold status is the recorded determination of a model's position relative to a capability threshold (not reached; cannot rule out; provisionally met; met), together with the date, evidentiary basis and determining party. This is element B5 of the source crosswalk. A cross-lab controlled vocabulary appears to be converging: GDM's "cannot rule out" ≈ OpenAI's "precautionarily treated as reached" ≈ Anthropic's "provisionally met" ≈ Meta's "provisionally rated" — the same precautionary intermediate state under four different names, observed across four labs.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. It is CASRAI's own reading of a published document. No lab, evaluator or regulator named here has declared, endorsed, or been consulted on this mapping. That will change only when an organisation files its own Mapping Declaration — see the non-endorsement policy.
OrganisationTheir term, as publishedMatchSource
Anthropic
Anthropic Risk Report, August 2026
Every frontier model since Opus 4 is treated as "provisionally meeting the CB-1 threshold" "to err on the side of caution" (§4.4.2). On AI R&D: "Neither RSP criterion is met" (§3.4).close
confidence: high
Anthropic Risk Report, August 2026
OpenAI
Preparedness Framework v2 / GPT-5.6 deployment safety report
SAG options: threshold crossed; threshold not crossed; recommend a deep dive (§3.3). GPT-5.6 designations: "High in Biological and Chemical", "High in Cybersecurity", "below High in AI Self-Improvement"; "these models should thus be precautionarily treated as High" (s.9, s.9.1.1).exact
confidence: high
OpenAI Preparedness Framework v2
Google DeepMind
Gemini 3.7 Flash FSF report / model card
"if we cannot rule out, based on the evidence and threat models we have, that a T/CCL has been reached, we designate the model as 'cannot rule out being at the T/CCL', and mitigate accordingly" (report p.4). Outcomes recorded as "No T/CCL reached" and "CBRN Uplift 1 CCL alert threshold reached" (pp.3, 12). Model card column "CCL reached?" with value "CCL not reached."exact
confidence: high
Gemini 3.7 Flash FSF report
xAI
Grok 4.6 card / Grok 4.20 card
"Grok 4.6 scores below the FAIF safety thresholds on dual-use knowledge, indicating limited actionable uplift for an already-trained actor." (§8, p.31). Grok 4.20: released "with safeguards appropriate for its capability threshold" (s.1.3), with the threshold left unidentified.close
confidence: medium
Grok 4.6 model card
Meta
Meta Advanced AI Scaling Framework v2
"Until evaluation on the complex suite of challenges is completed, any model meeting the simple-suite threshold is provisionally rated 'high' risk for the given deployment scenario" (§4.2.1, p.29). Assignment: "the Chief AI Officer and Director of Alignment and Risk will assign a risk threshold" (§2.1.2).close
confidence: high
Meta Advanced AI Scaling Framework v2
EU
EU GPAI Code of Practice, Safety and Security chapter
Framework must require "at least one systemic risk tier that has not been reached by the model" (Measure 4.1); the acceptance determination itself is a binary "acceptable" / "not determined to be acceptable" outcome per identified risk and overall (Measure 4.1(2)-(3), Measure 4.2), rather than a named "cannot rule out" intermediate status.close
confidence: high
EU GPAI Code of Practice, Safety and Security chapter
US Government (Executive Order 14409)
EO 14409
Developers may "engage the Federal Government to determine whether model(s) under development meet the designation of 'covered frontier model'" (Sec. 3(b)(i)) — a narrower, government-facing determination step rather than a full status vocabulary.narrow
confidence: medium
Executive Order 14409
METR
METR Common Elements
Thresholds "are compared to the results of model evaluations to determine whether they have been crossed."close
confidence: high
METR (common-elements)

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • Frontier Model Forum

    Crossing a threshold "signals entry into a new phase of heightened risk, where more rigorous risk assessments for this domain and stronger baseline safety and security measures are potentially warranted" (s3.1) — describes a consequence of status change, not a status vocabulary itself. Scored RL in the source document.

    FMF Risk Taxonomy and Thresholds

Gap

Controlled vocabulary emerging from the sources: not reached / alert threshold reached / cannot rule out (GDM) = precautionarily treated as reached (OpenAI) = provisionally met (Anthropic) = provisionally rated (Meta) / reached. The precautionary intermediate state is common to four labs under four different names — a strong candidate for a single NIKOLAI controlled-value ladder.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →