Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN6 · Mitigations and securityProposednikolai-v0.1

Safeguard

NIKOLAI proposes to define a Safeguard record as: a technical or procedural measure intended to reduce misuse or misalignment risk, identified by type (e.g. safety training, filter, monitoring, access control), the risk domain it targets, and the deployment scope it applies to. This is an unsourced NIKOLAI editorial synthesis; SB 53 leaves the parent term "mitigations" undefined and METR frames the closest analogue narrowly around access controls, so NIKOLAI's definition is deliberately broader than any single source.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. It is CASRAI's own reading of a published document. No lab, evaluator or regulator named here has declared, endorsed, or been consulted on this mapping. That will change only when an organisation files its own Mapping Declaration — see the non-endorsement policy.
OrganisationTheir term, as publishedMatchSource
Anthropic
Anthropic Risk Report, August 2026
Safeguards are described as a matrix of robustness levels (Level 1/2/3) × coverage levels × exemptions, mapped model-by-model in Table 4.5.A; ASL-2/ASL-3 terminology is explicitly retired for this purpose. "Required Safeguards": "a default set of required safeguards that are expected to bring risk down to acceptable levels"
"Required Safeguards" term is quoted by FMF {FMF31}.
exact
confidence: high
Anthropic Risk Report, August 2026
OpenAI
OpenAI Preparedness Framework v2 / GPT-5.6 deployment safety
"Safeguards Report" and "risk-specific safeguard guidelines"; "Preparedness Safeguards: safeguards deployed for the Tracked Categories rated High"exact
confidence: high
OpenAI Preparedness Framework v2
Google DeepMind
Frontier Safety Framework v3.1
"Deployment Mitigations: are safety measures we implement which are intended to counter the misuse or misaligned expression of critical capabilities in deployments." "Security Mitigations: ... intended to prevent the unauthorized modification or exfiltration of model weights"exact
confidence: high
Google DeepMind Frontier Safety Framework v3.1
xAI
Frontier AI Framework, 30 Jun 2026
"Safety training: 'Training our models to recognize and decline harmful requests.'" "System prompts". "Filters: 'Applying classifiers to verify safety when a model is queried regarding topics of CRBN [sic] risks'"
xAI's Frontier AI Framework (30 Jun 2026) carries PDF metadata reading "Privileged/Confidential DRAFT working FRAMEWORK DOC" with no xAI statement found disambiguating draft from final; treat this citation as provisional.
close
confidence: medium
xAI Frontier AI Framework, 30 June 2026 (draft-labeled PDF metadata)
Meta
Meta Advanced AI Scaling Framework v2
Table 1 measures ("Deploy with mitigations"); LlamaFirewall; Llama Defender programclose
confidence: medium
Meta Advanced AI Scaling Framework v2
EU
EU GPAI Code of Practice, Safety and Security Chapter
Commitment 5 "Safety mitigations" (training-data filtering, input/output monitoring and filtering, refusal fine-tuning, staged/gated access, "high-assurance quantitative safety guarantees", CoT-transparency techniques) and Commitment 6 "Security mitigations"; a model is "exempt from [Commitment 6] if the model's capabilities are inferior to the capabilities of at least one model for which the parameters are publicly available for download."exact
confidence: high
EU GPAI Code of Practice, Safety and Security Chapter
California SB 53
California SB 53
"Mitigations" undefined (framework topic (3))
SB 53 requires labs to describe mitigations but never defines the term itself — the gap is the finding, recorded rather than dropped.
none
confidence: low
California SB 53
METR
METR (metr.org)
"Model Deployment Mitigations: Access and model-level measures applied to prevent the unauthorized use of a model's dangerous capabilities."exact
confidence: high
METR
Safety Framework Cards (discovery)
Safety Framework Cards (SSRN 7061798, unread/paywalled — discovery sweep)
"mitigation commitments" [UV]
unverified:true — SSRN abstract page only; full text is account-gated and could not be read this pass, so the term's exact scope relative to Safeguard is unconfirmed.
none
confidence: low
Safety Framework Cards (SSRN 7061798), discovery sweep — unread/paywalled

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →