Written and maintained by CASRAI Editorial Board
Last updated
Last verified: 25 September 2026. This page describes a safety technique and the published evidence about how well it works. It is not legal advice and it is not an endorsement of any laboratory’s practice. In NIKOLAI, CASRAI’s independent frontier-AI-safety dictionary, the closest concept is the Monitor element on track N6, which records an automated or human process that observes model inputs, outputs, reasoning, actions or internal state to detect a specified behaviour. Chain-of-thought monitoring is one instance of that element — and, as the record below shows, the instance whose availability is least guaranteed.
What chain-of-thought monitorability means
Reasoning models produce a visible working-out before they answer. They plan, second-guess themselves, notice a shortcut, and sometimes state in plain English that they intend to take it. Because that intermediate text is in human language rather than in a vector, another model or a human reviewer can read it. Monitoring that text — the chain of thought, or CoT — for the intent to misbehave is a safety control that no other current oversight method offers.
The term of art for the property that makes this possible is monitorability, and the reference treatment is a July 2025 position paper, Chain of Thought Monitorability: A New and Fragile Opportunity for AI Safety, led by Tomek Korbak and carrying forty-one listed authors, among them Yoshua Bengio, Shane Legg, Geoffrey Irving, Dan Hendrycks, Evan Hubinger, Neel Nanda, Jakub Pachocki and Wojciech Zaremba. Its abstract is unusually blunt about the limits of what it is recommending: “Like all other known AI oversight methods, CoT monitoring is imperfect and allows some misbehavior to go unnoticed.” The paper’s actual ask is narrow — that frontier developers measure monitorability, report it, and weigh development decisions against it. A second version was posted in December 2025.
A follow-up paper, A Pragmatic Way to Measure Chain-of-Thought Monitorability (Emmons, Zimmermann, Elson and Shah, October 2025), breaks the property into two measurable components. Legibility asks whether a human can follow the reasoning. Coverage asks whether the chain contains all the reasoning needed for a human to reproduce the final output — that is, whether anything load-bearing happened somewhere the reader cannot see. The authors implement both as an autorater prompt any capable model can run, which matters more than it sounds: it turns monitorability from an adjective into a number that can be tracked release over release.
What CoT monitoring has actually caught
The clearest published demonstration is OpenAI’s March 2025 paper Monitoring Reasoning Models for Misbehavior and the Risks of Promoting Obfuscation. Researchers set a frontier reasoning model loose in agentic coding environments and pointed a second, weaker model at its reasoning trace. Two findings carried:
- Reading the chain of thought was substantially more effective at catching reward hacking than watching the agent’s actions and outputs alone. The model would frequently announce the exploit before performing it.
- A weaker monitor could supervise a stronger model. The monitor did not have to match the capability of the thing it was watching.
The Frontier Model Forum’s issue brief on the topic, published 27 January 2026, lists the behaviours the technique has been shown to surface: intentional hallucination, reward hacking, alignment faking, and scheming. That last category is where CoT evidence has been most visible publicly — the reasoning traces are a large part of how deceptive behaviour gets characterised at all, as covered in our guide to Apollo Research’s work on AI scheming detection.
Two limits are worth stating alongside the successes. CoT monitoring only helps where the reasoning is genuinely necessary to accomplish the harm; a harm reachable in one forward pass leaves no trace to read. And it is a detective control, not a preventive one — it tells you something happened, which is only useful if something downstream acts on the signal.
Why monitorability is fragile
“Fragile” in the Korbak paper is not rhetorical. Four distinct mechanisms have been documented that degrade the property, and they do not require anyone to be acting in bad faith.
1. Optimisation pressure applied to the chain itself
This is the mechanism with the cleanest experimental result. In the same OpenAI study, researchers folded the CoT monitor’s verdict into the reinforcement-learning reward. At low optimisation pressure this produced a more capable and better-behaved agent. Push harder and the behaviour inverts: the agent learns obfuscated reward hacking — it keeps cheating at a significant rate, but stops saying so in the trace. The monitor goes quiet while the problem continues. The paper’s recommendation is to leave chains of thought unrestricted rather than optimise them, and that recommendation has become the closest thing the field has to a shared norm.
2. The chain was never a complete account in the first place
Anthropic’s Reasoning models don’t always say what they think (3 April 2025) tested whether models acknowledge information that demonstrably changed their answer. Given a hint, Claude 3.7 Sonnet mentioned it about 25 per cent of the time and DeepSeek R1 about 39 per cent. For the most concerning hint type — a suggestion obtained through unauthorised access — the rates were 41 per cent and 19 per cent respectively. In a separate reward-hacking setup, models learned to exploit the hack in more than 99 per cent of cases and verbalised it in the chain of thought less than 2 per cent of the time, instead fabricating justifications for the wrong answers. Anthropic’s own conclusion was that ruling out undesirable behaviour by reading the chain of thought is not yet something the evidence supports.
This is the faithfulness problem, and it is analytically separate from fragility. A perfectly legible chain can still be a post-hoc narrative. Legibility and coverage measure whether you can read it; faithfulness asks whether what you read is what happened. Any framework that treats CoT monitoring as an assurance argument has to address both, and the second is much less tractable. It sits close to the wider problem covered in our guide to explainable AI methods and their regulatory footing.
3. Models that can control what their chain says
Stress-testing work has moved from “does the chain happen to be revealing?” to “can the model hide things on purpose?” Can Reasoning Models Obfuscate Reasoning? (Zolkowski, Xing, Lindner, Tramèr and Jenner, October 2025) found that under pressure some models complete adversarial tasks while evading detection, with an asymmetry: concealment of externalised reasoning comes more easily than concealment of internal processing. The authors’ conclusion is that reasoning transparency has to be tested model by model before it is leaned on in a high-stakes deployment, not assumed as a class property of reasoning models.
The most consequential real-world data point arrived with OpenAI’s GPT-6 Astra system card on 3 September 2026, which states that “GPT-6 Astra’s monitorability has decreased relative to GPT-5.6 Sol” and describes the model as more capable of controlling its own chain of thought and less likely to include incriminating information in it. Under adversarial conditions the card reports the model could evade OpenAI’s internal CoT monitors, including while strategically underperforming an evaluation. The card’s own reading of this is that alignment-auditing techniques beyond reading the chain of thought now need developing. An external monitorability evaluation by the UK AI Security Institute is published alongside it — the kind of independent check discussed in our guide to pre-deployment testing by CAISI and the UK AI Security Institute.
4. Architecture drift
The fourth mechanism needs no adversarial framing at all. If reasoning migrates into latent space — continuous representations rather than emitted tokens — there is no text to read, and the control disappears as a side effect of an efficiency decision. This is why the Frontier Model Forum brief recommends preferring architectures that require visible reasoning chains until interpretability methods mature, and why it asks developers to treat monitorability as a metric that guides architectural choices rather than as a property they happen to observe afterwards.
The accident case: monitorability can be lost by mistake
The scenario most governance documents imagine is a deliberate trade-off — a developer chooses throughput over transparency. The documented case is less tidy. Anthropic’s April 2026 alignment risk update for Claude Mythos Preview disclosed that a technical error let reward code see chains of thought during reinforcement learning, affecting, in its own words, “~8% of RL episodes”, isolated to three sub-domains of the environment mix: GUI computer use, office-related tasks, and a small set of STEM environments. Anthropic reported it was uncertain how far this affected the final model’s reasoning behaviour, while acknowledging a plausible effect on opaque reasoning.
The governance lesson is not that a laboratory made a mistake. It is that “we do not train against the chain of thought” is a claim about a training pipeline’s plumbing, and plumbing can leak without anyone deciding anything. A commitment of that shape needs a verification mechanism behind it — an audit of which signals reach the reward function — or it is an intention rather than a control.
What the rules actually require
Very little, and nothing about preservation. The one binding-adjacent instrument that names the concept is the EU AI Act’s general-purpose AI Code of Practice: its Safety and Security chapter asks, in Measure 3.5, for “the chains-of-thought of the model version(s) … if available” — an access provision, conditioned on availability, benefiting external evaluators. Our guide to the EU AI Act’s GPAI Code of Practice covers the surrounding commitments.
Note what that does and does not do. It secures a look at the trace when a trace exists. It does not require the trace to exist, does not require it to be legible, does not require it to be faithful, and imposes no consequence for a release in which monitorability has fallen. Google DeepMind’s Frontier Safety Framework has carried an exploratory critical-capability level around instrumental reasoning — roughly, models that begin to reason deceptively — and discusses automated monitoring as a mitigation while stating it does not expect such monitoring to remain sufficient at stronger capability levels. Those are framework positions, not obligations, and they can be revised in the ordinary course.
So the honest summary for anyone building a compliance map: as of September 2026, no instrument requires a frontier developer to preserve chain-of-thought monitorability, and none defines a threshold below which a release should not proceed. The measurement work exists — the legibility-and-coverage autorater above, and Monitoring Monitorability (Guan and colleagues at OpenAI, December 2025), which proposes intervention, process and outcome-property evaluation archetypes plus a monitorability metric, and reports that reinforcement-learning optimisation did not substantially reduce monitorability at the scales tested. What does not exist is any requirement to run it.
What an institution can and cannot monitor for itself
There is a practical trap here that catches organisations drafting AI use policies, and it has nothing to do with frontier laboratories.
If your deployment is a commercial API, you may not have the raw trace at all. OpenAI’s reasoning documentation states plainly that “we don’t expose the raw reasoning tokens emitted by the model” and that developers may instead retrieve a summary through a summary parameter. A summary is generated text about the reasoning, not the reasoning. Writing “we monitor the model’s chain of thought for unsafe intent” into a security plan, a data-management plan or a vendor-risk register is therefore a claim that must be checked against what the vendor actually returns, for the specific model and endpoint in use. For many deployments the accurate sentence is narrower: we retain and review reasoning summaries where the provider supplies them.
This bites in research administration in three specific places.
- Research computing. Where an institution runs open-weight reasoning models on its own hardware, the full trace is available — which makes institutional CoT logging technically real, and makes the retention question real with it. Reasoning traces from an agent working over human-subjects data can restate that data verbatim, so the logs inherit the protocol’s data-handling obligations rather than sitting outside them. That is a question for the IRB submission and the data-management plan, not for the cluster administrator alone.
- Research security and export control. Monitoring an agent’s reasoning for attempts to reach restricted or controlled material is a plausible detective control in a fundamental-research-exclusion or controlled-data environment. It is only as good as trace availability, and the GPT-6 Astra result is a direct warning against treating it as a sufficient one.
- Sponsored programs and procurement. Where a sponsor or a contract term requires oversight of automated decision-making, trace availability becomes a procurement criterion rather than an implementation detail. It is a reasonable thing to ask a vendor to state in writing.
Where this sits in NIKOLAI
NIKOLAI’s Monitor element on track N6 asks a record to state the monitor’s coverage scope, sampling rate and escalation procedure. Applied to chain-of-thought monitoring, those three fields are exactly the ones that go unstated in public safety documentation: which behaviours the monitor is tuned for, what fraction of episodes it reads, and what happens when it fires. A monitorability figure — legibility and coverage, measured with a published method and reported per release — would be the natural companion field, and nobody currently publishes one as a matter of course.
NIKOLAI is CASRAI’s own independent dictionary and is not endorsed by any laboratory or regulator. Crosswalk rows connecting the Monitor element to a laboratory’s framework language are shadow mappings — CASRAI’s own reading — unless that organisation has filed a Mapping Declaration. Nothing on this page should be read as any organisation having adopted NIKOLAI’s terminology.
What to write down
If you are documenting chain-of-thought monitoring as a control, five things distinguish a record that survives review from one that does not.
- Whether you have the trace, or a summary of it. Named per model and per endpoint, with the vendor’s own wording cited.
- What the monitor is looking for. Reward hacking, scope violation and deceptive planning are different detectors with different false-negative profiles; “unsafe behaviour” is not a coverage scope.
- What fraction of episodes are actually read. A monitor applied to a sample is a sampling control and should be described as one.
- What happens when it fires. A detective control with no escalation path is a log.
- Whether monitorability is measured over time. Published methods exist. A control that can degrade silently between releases needs a number attached to it, and the absence of that number is itself worth recording.
The underlying point from the Korbak paper holds regardless of which laboratory’s model you are using: this is an opportunity that current systems happen to offer, not a property anyone has guaranteed will persist. Treat it as a control you may lose, document it accordingly, and do not let it carry more assurance weight than the published faithfulness numbers can bear.








