Source of record
Where this definition comes from
Anthropic Risk Report, August 2026, §4.5, fn 59
“Safeguards are described as a matrix of robustness levels (Level 1/2/3) × coverage levels × exemptions, mapped model-by-model in Table 4.5.A; ASL-2/ASL-3 terminology is explicitly retired for this purpose”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfGoogle DeepMind Frontier Safety Framework v3.1, glossary
“Deployment Mitigations: are safety measures we implement which are intended to counter the misuse or misaligned expression of critical capabilities in deployments.”
https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3-1.pdfMETR, (element level)
“Model Deployment Mitigations: Access and model-level measures applied to prevent the unauthorized use of a model's dangerous capabilities.”
https://metr.org/
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report, August 2026 | “Safeguards are described as a matrix of robustness levels (Level 1/2/3) × coverage levels × exemptions, mapped model-by-model in Table 4.5.A; ASL-2/ASL-3 terminology is explicitly retired for this purpose. "Required Safeguards": "a default set of required safeguards that are expected to bring risk down to acceptable levels"” "Required Safeguards" term is quoted by FMF {FMF31}. | exact confidence: high | Anthropic Risk Report, August 2026 |
| OpenAI OpenAI Preparedness Framework v2 / GPT-5.6 deployment safety | “"Safeguards Report" and "risk-specific safeguard guidelines"; "Preparedness Safeguards: safeguards deployed for the Tracked Categories rated High"” | exact confidence: high | OpenAI Preparedness Framework v2 |
| Google DeepMind Frontier Safety Framework v3.1 | “"Deployment Mitigations: are safety measures we implement which are intended to counter the misuse or misaligned expression of critical capabilities in deployments." "Security Mitigations: ... intended to prevent the unauthorized modification or exfiltration of model weights"” | exact confidence: high | Google DeepMind Frontier Safety Framework v3.1 |
| xAI Frontier AI Framework, 30 Jun 2026 | “"Safety training: 'Training our models to recognize and decline harmful requests.'" "System prompts". "Filters: 'Applying classifiers to verify safety when a model is queried regarding topics of CRBN [sic] risks'"” xAI's Frontier AI Framework (30 Jun 2026) carries PDF metadata reading "Privileged/Confidential DRAFT working FRAMEWORK DOC" with no xAI statement found disambiguating draft from final; treat this citation as provisional. | close confidence: medium | xAI Frontier AI Framework, 30 June 2026 (draft-labeled PDF metadata) |
| Meta Meta Advanced AI Scaling Framework v2 | “Table 1 measures ("Deploy with mitigations"); LlamaFirewall; Llama Defender program” | close confidence: medium | Meta Advanced AI Scaling Framework v2 |
| EU EU GPAI Code of Practice, Safety and Security Chapter | “Commitment 5 "Safety mitigations" (training-data filtering, input/output monitoring and filtering, refusal fine-tuning, staged/gated access, "high-assurance quantitative safety guarantees", CoT-transparency techniques) and Commitment 6 "Security mitigations"; a model is "exempt from [Commitment 6] if the model's capabilities are inferior to the capabilities of at least one model for which the parameters are publicly available for download."” | exact confidence: high | EU GPAI Code of Practice, Safety and Security Chapter |
| California SB 53 California SB 53 | “"Mitigations" undefined (framework topic (3))” SB 53 requires labs to describe mitigations but never defines the term itself — the gap is the finding, recorded rather than dropped. | none confidence: low | California SB 53 |
| METR METR (metr.org) | “"Model Deployment Mitigations: Access and model-level measures applied to prevent the unauthorized use of a model's dangerous capabilities."” | exact confidence: high | METR |
| Safety Framework Cards (discovery) Safety Framework Cards (SSRN 7061798, unread/paywalled — discovery sweep) | “"mitigation commitments" [UV]” unverified:true — SSRN abstract page only; full text is account-gated and could not be read this pass, so the term's exact scope relative to Safeguard is unconfirmed. | none confidence: low | Safety Framework Cards (SSRN 7061798), discovery sweep — unread/paywalled |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- Frontier Model Forum
"Safeguards: Improving defenses against known attack methods, such as adversarial inputs or data poisoning, for safety and security purposes only" (Table 3, a use category) — related, not a mapping (RL).
Frontier Model Forum, Information Sharing / Incident Reporting Issue Brief







