Source of record
Where this definition comes from
Anthropic Risk Report, August 2026, §2.23.1.1
“an Opus 4.6 completion-only monitor (no subsampling for the most common agent scaffolds; 5% for tool-marked completions), an Opus 4.8 prompt+completion monitor on a 25% subsample of impactful-tool-call traffic ... surface ~50 trajectories per week for manual review by AI control researchers”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfOpenAI Preparedness Framework v2, Table 5
“Monitor AI supervises agent actions”
https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdfEU GPAI Code of Practice, Safety and Security Chapter, Measure 3.5
“will conduct appropriate post-market monitoring to gather information relevant to assessing whether the systemic risk could be determined to not be acceptable ... and to inform whether a Model Report update is necessary”
https://ec.europa.eu/newsroom/dae/redirection/document/118119
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report, August 2026 / Training a Misaligned Reward Seeker | “an Opus 4.6 completion-only monitor (no subsampling for the most common agent scaffolds; 5% for tool-marked completions), an Opus 4.8 prompt+completion monitor on a 25% subsample of impactful-tool-call traffic ... surface ~50 trajectories per week for manual review by AI control researchers; "we monitor the Chain-of-Thought to detect and measure reward hacking, but importantly do not reward or penalize the model based on it, in order to preserve monitorability"” Term also draws on {RSEEK} ("Training a Misaligned Reward Seeker", 16 Sep 2026 correction pass) for the CoT-monitoring/severity-flagging quote. | exact confidence: high | Anthropic Risk Report, August 2026 |
| OpenAI OpenAI Preparedness Framework v2 / Pacing Model Development (Cyber), Aug 2026 | “"Monitor AI supervises agent actions"; "Monitoring (safeguard): 'which detects and allows us to respond to concerning behavior'"; "Activation classifiers: 'detectors that inspect a model's internal activity'"; "Automated investigators"” Also cites {OPACE} (Aug 2026 post) for activation classifiers / automated investigators language. | exact confidence: high | OpenAI Preparedness Framework v2 |
| Google DeepMind Frontier Safety Framework v3.1 | “Deployment mitigations include "input/output/chain-of-thought monitoring and analysis" (glossary); safeguards "such as chain-of-thought monitoring on high-risk deployments"” | close confidence: medium | Google DeepMind Frontier Safety Framework v3.1 |
| xAI Frontier AI Framework, 30 Jun 2026 | “Incident detection "Monitoring and alerting of public comments from the X platform"; security "real-time monitoring"” xAI's Frontier AI Framework (30 Jun 2026) carries PDF metadata reading "Privileged/Confidential DRAFT working FRAMEWORK DOC" with no xAI statement found disambiguating draft from final; treat this citation as provisional. | narrow confidence: medium | xAI Frontier AI Framework, 30 June 2026 (draft-labeled PDF metadata) |
| Meta Meta Advanced AI Scaling Framework v2 | “LlamaFirewall "chain-of-thought auditing" and input classifiers” | narrow confidence: medium | Meta Advanced AI Scaling Framework v2 |
| EU EU GPAI Code of Practice, Safety and Security Chapter | “Measure 3.5 "Post-market monitoring": Signatories "will conduct appropriate post-market monitoring to gather information relevant to assessing whether the systemic risk could be determined to not be acceptable ... and to inform whether a Model Report update is necessary", via methods including end-user feedback, incident reporting forms, bug bounties, "monitoring software repositories, known malware, public forums, and/or social media", and "monitoring aspects of models that are relevant for assessing and mitigating systemic risk and are not transparent to third parties, e.g. hidden chains-of-thought"” Source notes this confirms xAI's chapter-derived s.2.2(5) wording; the linked {FAIF26} carries the draft/final metadata caveat above. | exact confidence: high | EU GPAI Code of Practice, Safety and Security Chapter |
| Frontier Model Forum FMF Information Sharing / Incident Reporting Issue Brief | “"Monitoring and Detection Systems: Enhancing systems that detect anomalous behavior, unauthorized access, or potential misuse, for safety and security purposes only" (Table 3)” | close confidence: medium | Frontier Model Forum, Information Sharing / Incident Reporting Issue Brief |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- SB 53 (California)
Framework topic "(10) ... including risks resulting from a frontier model circumventing oversight mechanisms" (22757.12(a)) — related to monitoring but a pointer, not a mapping (RL).
California SB 53 - UK AISI / Google DeepMind
"CoT monitoring helps us understand how an AI system produces its answers, complementing interpretability research." — related, not a direct monitor-record mapping (RL).
DeepMind, Deepening Our Partnership with UK AI Security Institute







