Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN6 · Mitigations and securityProposednikolai-v0.1

Monitor

NIKOLAI proposes to define a Monitor record as: an automated or human process that observes model inputs, outputs, reasoning, actions, or internal state to detect a specified behavior, carrying a stated coverage scope, sampling rate, and escalation path when a detection fires. This is an unsourced NIKOLAI editorial synthesis of how Anthropic, OpenAI, and Google DeepMind each independently operate chain-of-thought and output monitors as a distinct mitigation category from access-control safeguards.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. It is CASRAI's own reading of a published document. No lab, evaluator or regulator named here has declared, endorsed, or been consulted on this mapping. That will change only when an organisation files its own Mapping Declaration — see the non-endorsement policy.
OrganisationTheir term, as publishedMatchSource
Anthropic
Anthropic Risk Report, August 2026 / Training a Misaligned Reward Seeker
an Opus 4.6 completion-only monitor (no subsampling for the most common agent scaffolds; 5% for tool-marked completions), an Opus 4.8 prompt+completion monitor on a 25% subsample of impactful-tool-call traffic ... surface ~50 trajectories per week for manual review by AI control researchers; "we monitor the Chain-of-Thought to detect and measure reward hacking, but importantly do not reward or penalize the model based on it, in order to preserve monitorability"
Term also draws on {RSEEK} ("Training a Misaligned Reward Seeker", 16 Sep 2026 correction pass) for the CoT-monitoring/severity-flagging quote.
exact
confidence: high
Anthropic Risk Report, August 2026
OpenAI
OpenAI Preparedness Framework v2 / Pacing Model Development (Cyber), Aug 2026
"Monitor AI supervises agent actions"; "Monitoring (safeguard): 'which detects and allows us to respond to concerning behavior'"; "Activation classifiers: 'detectors that inspect a model's internal activity'"; "Automated investigators"
Also cites {OPACE} (Aug 2026 post) for activation classifiers / automated investigators language.
exact
confidence: high
OpenAI Preparedness Framework v2
Google DeepMind
Frontier Safety Framework v3.1
Deployment mitigations include "input/output/chain-of-thought monitoring and analysis" (glossary); safeguards "such as chain-of-thought monitoring on high-risk deployments"close
confidence: medium
Google DeepMind Frontier Safety Framework v3.1
xAI
Frontier AI Framework, 30 Jun 2026
Incident detection "Monitoring and alerting of public comments from the X platform"; security "real-time monitoring"
xAI's Frontier AI Framework (30 Jun 2026) carries PDF metadata reading "Privileged/Confidential DRAFT working FRAMEWORK DOC" with no xAI statement found disambiguating draft from final; treat this citation as provisional.
narrow
confidence: medium
xAI Frontier AI Framework, 30 June 2026 (draft-labeled PDF metadata)
Meta
Meta Advanced AI Scaling Framework v2
LlamaFirewall "chain-of-thought auditing" and input classifiersnarrow
confidence: medium
Meta Advanced AI Scaling Framework v2
EU
EU GPAI Code of Practice, Safety and Security Chapter
Measure 3.5 "Post-market monitoring": Signatories "will conduct appropriate post-market monitoring to gather information relevant to assessing whether the systemic risk could be determined to not be acceptable ... and to inform whether a Model Report update is necessary", via methods including end-user feedback, incident reporting forms, bug bounties, "monitoring software repositories, known malware, public forums, and/or social media", and "monitoring aspects of models that are relevant for assessing and mitigating systemic risk and are not transparent to third parties, e.g. hidden chains-of-thought"
Source notes this confirms xAI's chapter-derived s.2.2(5) wording; the linked {FAIF26} carries the draft/final metadata caveat above.
exact
confidence: high
EU GPAI Code of Practice, Safety and Security Chapter
Frontier Model Forum
FMF Information Sharing / Incident Reporting Issue Brief
"Monitoring and Detection Systems: Enhancing systems that detect anomalous behavior, unauthorized access, or potential misuse, for safety and security purposes only" (Table 3)close
confidence: medium
Frontier Model Forum, Information Sharing / Incident Reporting Issue Brief

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • SB 53 (California)

    Framework topic "(10) ... including risks resulting from a frontier model circumventing oversight mechanisms" (22757.12(a)) — related to monitoring but a pointer, not a mapping (RL).

    California SB 53
  • UK AISI / Google DeepMind

    "CoT monitoring helps us understand how an AI system produces its answers, complementing interpretability research." — related, not a direct monitor-record mapping (RL).

    DeepMind, Deepening Our Partnership with UK AI Security Institute

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →