Skip to main content
v2026.11,858 entries · CC-BY 4.0
Track N6

Mitigations and security

Monitor, safeguard, robustness level, coverage level, exemption, security level, security control, mitigation change, internal deployment and internal-use risk.

Once a threshold is crossed, a framework has to say what actually changes: which safeguard, at what robustness level, covering what -- plus the security controls, including model-weight security, that keep the safeguard itself from being bypassed. This track defines safeguard, security level, security control, robustness level, and monitor, as specified by Anthropic, OpenAI, Google DeepMind, xAI, Meta, Amazon, Microsoft, and G42.

  • Internal deployment and internal-use risk
    Proposedrecord-type

    A proposed record type for a developer's own internal use of a model, including agentic research and training-time use, kept distinct from public deployment.

  • Mitigation change
    Proposedrecord-type

    A proposed record for a logged change to a safeguard or security control, capturing its date, scope, reason, and effect on reassessment.

  • Security control
    Proposedrecord-type

    A proposed record for one security measure, such as access approval or weight encryption, named and mapped to an external control catalogue where possible.

  • Security level
    Proposedcontrolled-value

    A proposed graded statement of a developer's model-weight and infrastructure security posture, indexed to a named external scale where one is adopted.

  • Exemption (reduced-safeguard or trusted access)
    Proposedrecord-type

    A proposed record for an authorized arrangement where named users get a model with safeguards reduced or removed, plus its vetting and revocation terms.

  • Coverage level
    Proposedcontrolled-value

    A proposed graded scale describing what content or behavior a safeguard is designed to catch, kept separate from how hard it is to circumvent.

  • Robustness level
    Proposedcontrolled-value

    A proposed graded scale for how hard a safeguard is to circumvent, paired with the assumed attack model and the evidence behind the grade.

  • Safeguard
    Proposedrecord-type

    A proposed record for a technical or procedural measure meant to reduce misuse or misalignment risk, identified by type, target risk, and deployment scope.

  • Monitor
    Proposedrecord-type

    A proposed record for a process that watches model inputs, outputs, or actions to detect a behavior, with its coverage, sampling rate, and escalation path.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →