Source of record
Where this definition comes from
Anthropic Risk Report, August 2026, §2.1, §2.19
“Alignment risk "explicitly raised from 'very low' to reflect increased uncertainty after 'recent incident disclosures related to model behavior in cybersecurity evaluations'" (§2.1, §2.19).”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfGemini 3.7 Flash FSF Report, p.4
“"Safety margin: takes into account different sources of uncertainty, including the limits of our threat modeling and evaluations" (p.4); "cannot rule out being at the T/CCL".”
https://storage.googleapis.com/deepmind-media/gemini/gemini_3-7_flash_fsf_report.pdf
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report, August 2026 | “Alignment risk "explicitly raised from 'very low' to reflect increased uncertainty after 'recent incident disclosures related to model behavior in cybersecurity evaluations'" (§2.1, §2.19).” | exact confidence: high | Anthropic Risk Report, August 2026 |
| OpenAI OpenAI Frontier Governance Framework / Preparedness Framework v2 | “"out of an abundance of caution we have treated models as crossing a capability threshold in circumstances where we are unable to rule out that a new threshold had been reached, even in the absence of direct evidence that it has occurred" (FGF §2.3). Sandbagging response: "use a conservative upper bound of the model's non-sandbagged evaluation results" (PF Table 2).” | close confidence: high | OpenAI Frontier Governance Framework |
| Google DeepMind Gemini 3.7 Flash FSF Report | “"Safety margin: takes into account different sources of uncertainty, including the limits of our threat modeling and evaluations" (p.4); "cannot rule out being at the T/CCL".” | exact confidence: high | Gemini 3.7 Flash FSF Report |
| xAI xAI Frontier AI Framework, 30 Jun 2026 | “"incorporating a margin of security" (s.2.3); not quantified.” This row cites xAI's Frontier AI Framework (30 Jun 2026, {FAIF26}), whose PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC" with no xAI statement found disambiguating draft vs. final status. Treat as citing a document of unconfirmed draft/final status. | none confidence: medium | xAI Frontier AI Framework, 30 Jun 2026 |
| Meta Meta Advanced AI Scaling Framework v2 | “"We conduct risk assessments and assign risk thresholds with maximum elicitation in mind, capturing the upper bound of risk" (§2.2.1); "provisionally rated 'high'" (§4.2.1).” | close confidence: high | Meta Advanced AI Scaling Framework v2 |
| EU EU GPAI Code of Practice, Safety and Security Chapter | “Measure 4.1: acceptance determination "incorporat[es] a safety margin" that must "(1) be appropriate for the systemic risk; and (2) take into account potential limitations, changes, and uncertainties of: (a) systemic risk sources (e.g. capability improvements after the time of assessment); (b) systemic risk assessments (e.g. under-elicitation of model evaluations or historical accuracy of similar assessments); and (c) the effectiveness of safety and security mitigations" -- the only source that itemises what a safety margin must account for, matching xAI's chapter-derived "margin of security" language (fn.1-2).” This row's source citation also references xAI's Frontier AI Framework, 30 Jun 2026 ({FAIF26}), for the correlated 'margin of security' language it echoes. That xAI document's PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC" with no xAI statement found disambiguating draft vs. final status; the EU citation itself (EUSSC, the official 43-page chapter) is not affected, but any inference drawn about xAI's own framework from this correlation should carry the same draft-status caveat. | exact confidence: high | EU GPAI Code of Practice, Safety and Security Chapter |







