Skip to main content
v2026.11,858 entries · CC-BY 4.0

When AI Reviews AI: Inside NIKOLAI’s AI-Model-Review Element

Anthropic’s August 2026 Risk Report names the exact model instance that reviewed a draft of its own alignment section — Claude Mythos 5, with named internal access, in 24 minutes. That is not human-evaluator independence; it’s an AI system acting as an assurance actor. NIKOLAI’s AI-Model Review element, in CASRAI’s own independent dictionary, is built to record exactly that — and to name the human still accountable for what the AI found.

Written and maintained by CASRAI Editorial Board

Last updated

Last verified: September 20, 2026. In Section 2.20 of its August 2026 Risk Report, Anthropic did something no frontier lab’s safety report had done quite this explicitly before: it named the exact model instance that reviewed a near-final draft of the report’s own alignment section, what internal access that instance was given, and how long the task took. The reviewer was Claude Mythos 5. Its access included Anthropic’s internal Slack channels, a large swath of internal documents, the internal codebase, and the ability to direct subagents to investigate specific questions. The task — reading the rest of Section 2, comparing it against the internal record, and producing a signed, publishable critique — took Claude Mythos 5 24 minutes.

That is not a hypothetical case study. It is a citable line item, in a document a frontier lab put its own name on, describing an AI system performing what the lab itself calls “an experimental source of additional assurance.” Human-evaluator independence — who checks a lab’s safety claims, and whether that checker has a conflict of interest — is a well-covered question (CASRAI’s own guide on evaluator independence covers it). This is a different question: what happens when the checker isn’t human at all? NIKOLAI, CASRAI’s own frontier-AI-safety dictionary, has an element for exactly that: AI-Model Review. This guide is a deep dive on that element, anchored to the one case, so far, where a lab has published enough detail to actually fill it in.

What Anthropic Actually Disclosed

Section 2.20 of the Risk Report, titled “Claude’s review of this risk assessment,” describes prompting an instance of Claude Mythos 5 to review a near-final draft of the rest of Section 2 (the section covering alignment risk), and to produce a critique suitable for publication alongside the report itself. The prompt Anthropic used is reproduced in the report verbatim; so is the model’s full response. A few details matter for what follows:

  • Access level. The report lists what Mythos 5’s review scaffold gave it: access to “many of our internal Anthropic Slack channels (encompassing the vast majority of the internal discussion of this report),” “a large swath of relevant internal documents,” “our internal codebase,” and the ability to “direct targeted subagents to investigate particular questions as desired.”
  • Task duration. A footnote to the transcript (labeled “Transcript 2.20.A” in the report) states plainly: “In practice, Claude took 24 minutes to produce this review.” The report is careful to note that Claude’s own reference to “bounded time” in its written response wasn’t a constraint Anthropic imposed — the model volunteered it, later clarifying it just meant to signal this wasn’t a multi-day human-team review.
  • Output disposition. Claude’s review raised three specific criticisms (about Section 2.9.4’s treatment of training-data contamination, a fully-redacted incident in Section 2.23.1.2, and how Section 2.19’s risk-level increase was attributed). Anthropic’s report states it added caveats to Section 2.9.4 as a direct result of the feedback, and that it regards the other two criticisms as reasonable but chose not to act on them in this cycle.
  • Accountable owner. The disposition decision — what to do with an AI reviewer’s findings — stayed with Anthropic’s human authors. The report is explicit that publishing the review was a deliberate choice, and that the review’s own text, not a paraphrase, is what got published.

Every one of those is a field NIKOLAI’s AI-Model Review element asks a record like this to capture: system identity, task, access granted, supervision applied, output disposition, and the accountable human owner of that disposition. Anthropic didn’t fill out a NIKOLAI record — NIKOLAI isn’t a form anyone submits to — but Section 2.20 is, functionally, the first public disclosure detailed enough that CASRAI could map it against that structure at all.

The NIKOLAI Element: AI-Model Review

AI-Model Review sits in NIKOLAI’s N8 track (Transparency and Review), alongside evaluator independence, evaluator access attestation, publication-rights clauses, external review, and redaction. As of this writing the element carries status “Proposed” in the current NIKOLAI release. Its definition, verified directly against the live element page, is: a record that an AI system performed an assurance task — review, monitoring, grading, red-teaming, or analysis — capturing the system’s identity, the task, the access it was granted, the supervision applied, the disposition of its output, and the accountable human owner of that disposition. CASRAI’s own note on the element is explicit that this is distinct from human-authorship credit frameworks like CRediT — a distinction this guide comes back to below.

NIKOLAI’s crosswalk table for AI-Model Review currently lists six organizations. Every row is what CASRAI calls a shadow mapping: CASRAI’s own independent reading of what each organization has published, not a mapping any of these organizations has reviewed, confirmed, or endorsed. That caveat is on the live element page itself, and it’s worth repeating here rather than softening it:

  • Anthropic — Claude model reviews of internal documents, cited directly to the “Claude took 24 minutes to produce this review” language in the Risk Report. Match type: Equivalent. This is the only row built on a primary-source quote this specific.
  • OpenAI — mapped, at Close confidence, to OpenAI’s own description of a “Monitor AI” that supervises agent actions, and to “GPT-Red” as a red-teaming model.
  • Google DeepMind — mapped, at Close confidence, to DeepMind’s “Investigator agent” and “Prompted Classifiers” used for malicious-intent labeling.
  • xAI — mapped, at Close confidence, to what NIKOLAI describes as an “automated alignment audit” performed via model grading.
  • Meta — mapped, at Narrow confidence, to LlamaFirewall’s chain-of-thought auditing.
  • METR — mapped, at Close confidence, to GPT-5.6 Sol analysis agents used for incident investigation.

The confidence and match-type labels are doing real work here, not decoration. “Equivalent” on the Anthropic row means CASRAI judges Anthropic’s own described practice to line up closely with what the element defines. “Close” and “Narrow” on the other five mean the underlying practice is related but not a clean match — a monitoring system, a classifier, or an audit agent isn’t necessarily performing the same kind of bounded, disclosed, single-instance review Section 2.20 describes. None of these six organizations has filed a NIKOLAI Mapping Declaration confirming any of this. If one does, that row moves from shadow mapping to declared and this guide will be updated to say so.

The Accountability Gap This Element Is Built to Name

CRediT (ANSI/NISO Z39.104-2022) is the standard CASRAI stewards for describing who did what on a piece of scholarly or reference work — conceptualization, writing, validation, and eleven other roles. It is, deliberately, a taxonomy for human contributors. CASRAI’s own guide on the question is direct about this: CRediT roles do not apply to AI tools, because CRediT presupposes an accountability relationship — someone who can answer for a contribution — that current frontier AI systems don’t hold. What CRediT can capture is a human’s verification or oversight work when that human used an AI tool as part of their own contribution; CASRAI’s separate guide on disclosing AI assistance in CRediT statements covers that mechanism.

That leaves a real gap, and Section 2.20 sits directly inside it. Claude Mythos 5 performed a bounded, disclosed, consequential task — a safety-relevant review that changed what Anthropic published — and there is no CRediT role, no authorship credit, no contributor-taxonomy entry that describes what it did. It isn’t an author. It isn’t a human contributor whose oversight work can be logged. It’s a non-human system that did a specific, time-boxed job with defined access, and whose output a human then had to decide what to do with. CRediT was never built to answer “what did the AI actually do, and who signed off on it,” because CRediT’s whole design starts from the premise that the contributor is a person who can be held accountable. AI-Model Review starts from the opposite premise: the system doing the work can’t be accountable, so the record has to capture who is — explicitly, as one of its named fields, not as an afterthought.

This is why NIKOLAI treats AI-Model Review as its own element rather than folding it into evaluator independence (which is about whether a human or institutional evaluator has a conflict of interest) or into CRediT-adjacent AI-disclosure guidance (which is about a human author’s use of an AI tool). Section 2.20 isn’t a disclosure that a human used AI assistance while writing the report. It’s a disclosure that an AI system performed the assurance task itself, on a section that a human then chose whether to revise. That’s a third category, and until Section 2.20, there wasn’t a public, primary-source example specific enough to test whether a structured record for it was even worth having.

Why This Is CASRAI’s Own Framing, Not Anthropic’s

It bears repeating in its own section, because it’s easy to blur: NIKOLAI is CASRAI’s own project, built and maintained independently, and every crosswalk mapping on the AI-Model Review element — including the Anthropic row — is a shadow mapping unless and until Anthropic files a Mapping Declaration saying otherwise. Anthropic published Section 2.20 as part of its own Risk Report, for its own reasons, using its own vocabulary; nothing in the report references NIKOLAI, and Anthropic has not endorsed, reviewed, or been consulted on how CASRAI classifies that disclosure. The “Claude Mythos 5” name, the “24 minutes” figure, and the access-level details cited above are quoted directly from Anthropic’s own published Risk Report, not synthesized or inferred by NIKOLAI — that’s exactly why the Anthropic row is the one row on the element marked “Equivalent” at “High” confidence rather than “Close” or “Narrow.” The other five rows are CASRAI’s own reading of publicly available material about OpenAI, Google DeepMind, xAI, Meta, and METR, and carry the lower-confidence labels that reflect how much interpretation went into each one.

Frequently Asked Questions

Is “Claude Mythos 5” a real Anthropic model, and is this detail really in Anthropic’s Risk Report?

Yes. Section 2.20 of Anthropic’s own August 2026 Risk Report names Claude Mythos 5 as the model instance that reviewed a draft of the report’s alignment section, describes its internal access, and states in a footnote that the task took 24 minutes. This guide quotes that section directly rather than relying on NIKOLAI’s synthesis for this specific detail.

Is NIKOLAI’s AI-Model Review element an official or endorsed framework?

No. It’s a proposed element in CASRAI’s own, independent, unendorsed NIKOLAI dictionary. None of the six organizations on its crosswalk table — Anthropic included — has reviewed, confirmed, or endorsed how NIKOLAI classifies their practice. Every row is a shadow mapping unless a Mapping Declaration says otherwise.

How is this different from evaluator independence?

Evaluator independence (also on NIKOLAI’s N8 track) is about whether a human or institutional evaluator reviewing a lab’s safety claims has a financial or personal conflict of interest. AI-Model Review is about an AI system itself performing the review, monitoring, or red-teaming task — a different actor, a different set of questions (access granted, task duration, output disposition), and a different accountability problem, since the reviewer here can’t hold a conflict of interest in the human sense at all.

Could CRediT just add an “AI reviewer” role instead?

CASRAI’s position, reflected in its own CRediT guidance, is no: CRediT’s roles presuppose an accountability relationship that current AI systems don’t hold, so adding an AI-specific role would break the taxonomy’s core logic rather than extend it. That’s precisely the gap NIKOLAI’s AI-Model Review element is built to fill with a separate, purpose-built record instead.

Related Reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about When AI Reviews AI: Inside NIKOLAI’s AI-Model-Review Element

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →