Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN6 · Mitigations and securityProposednikolai-v0.1

Exemption (reduced-safeguard or trusted access)

NIKOLAI proposes an Exemption record as: an authorised arrangement under which named users or organisations receive a model with safeguards reduced or removed, carrying fields for vetting criteria, monitoring terms, attestation, and revocation conditions. This is an unsourced NIKOLAI editorial synthesis. Both 2026 evaluation-environment incidents in the source corpus involved models running under reduced safeguards inside an evaluation arrangement, and SB 53's carve-out excludes such access from its statutory definition of "deployment" — which makes an exemption record load-bearing for incident and governance tracking rather than administrative detail.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. It is CASRAI's own reading of a published document. No lab, evaluator or regulator named here has declared, endorsed, or been consulted on this mapping. That will change only when an organisation files its own Mapping Declaration — see the non-endorsement policy.
OrganisationTheir term, as publishedMatchSource
Anthropic
Anthropic Risk Report, August 2026 / Investigating Incidents (Cybersecurity Evals)
"Exemptions — programmes by which 'some trusted and vetted users [may] access our models with reduced or no CB-risk-based blocking classifiers'"; "Helpful-only models: 'variants of our production models — models trained not to refuse on potentially harmful user requests'" governed by a "Helpful-Only Model Access Policy"exact
confidence: high
Anthropic Risk Report, August 2026
OpenAI
GPT-5.6 system card / Path to Astra / Preparedness Framework v2
"Trusted Access for Cyber (TAC): 'an identity-gated access pathway that provides higher-risk dual-use cyber capabilities to enterprise customers, verified defenders, and other legitimate users'"; "Trusted Access for Biology Research"; trusted access programmes "do not remove monitoring or permit the highest-risk categories of assistance"; PF claim "Trust-based Access"
Citation spans three OpenAI documents ({G56}, {ASTRA}, {PF}); primary term quoted from the GPT-5.6 system card.
close
confidence: medium
OpenAI GPT-5.6 System Card / Deployment Safety
xAI
Frontier AI Framework, 30 Jun 2026 / 31 Dec 2025
"the full functionality of our models may be available to only a limited set of trusted parties, partners, and government agencies" (Jun 2026); "we may selectively allow xAI's models to respond to such requests from some vetted, highly trusted users (such as trusted third-party safety auditors or large enterprise customers under contract)" (Dec 2025)
xAI's 30 Jun 2026 Frontier AI Framework carries PDF metadata reading "Privileged/Confidential DRAFT working FRAMEWORK DOC" with no xAI statement found disambiguating draft from final; treat the {FAIF26}-sourced portion of this citation as provisional.
close
confidence: medium
xAI Frontier AI Framework, 30 June 2026 (draft-labeled PDF metadata)
Meta
Meta Advanced AI Scaling Framework v2
"We believe that there are numerous potential benefits of allowing trusted researchers working on scientific research to access the chemical and biological capabilities of highly capable AI models with reduced mitigations via structured access controls."close
confidence: medium
Meta Advanced AI Scaling Framework v2
Anthropic
Anthropic, Improving Alignment & Security Efforts
Anthropic asked "every organization that tests pre-release models with reduced cyber safeguards to commit to a set of best practices" (sandbox and network isolation, pre-engagement validation, explicit scope-setting, real-time monitoring)
Listed as "Other" in the source evidence table but is itself an Anthropic document, distinct from the {RR} row above.
close
confidence: medium
Anthropic, Improving Alignment & Security Efforts

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • Google DeepMind

    "A critical capability assessment may not be conducted for low-risk external deployments (e.g. to a small number of trusted testers) if the appropriate governance function determines the residual risk of such deployments to be acceptable"; "Out-of-model safeguards" disabled for red teaming — related, not a mapping (RL).

    Google DeepMind Frontier Safety Framework v3.1
  • SB 53 (California)

    "'Deploy' does not include making a frontier model available to a third party for the primary purpose of developing or evaluating the frontier model." — related statutory carve-out, not a record mapping (RL); load-bearing for why an exemption record matters (see operational_definition).

    California SB 53
  • US Government (EO 14409)

    "trusted partners that will have early access to covered frontier models" — related, not a mapping (RL).

    Executive Order 14409
  • Frontier Model Forum

    Access item "version with limited safety training" — related, not a mapping (RL).

    Frontier Model Forum, Third-Party Assessments

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →