Source of record
Where this definition comes from
Anthropic Risk Report, August 2026, §2.20
“we prompted an instance of Claude Mythos 5 to review a near-final draft of Section 2 of this report ... Readers should weigh my position honestly, as I do: I am a Claude model reviewing Anthropic's assessment of Claude models ... In practice, Claude took 24 minutes to produce this review.”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfClaude's Constitution, Acknowledgements
“Several Claude models provided feedback on drafts. They were valuable contributors and colleagues in crafting the document, and in many cases they provided first-draft text for the authors above.”
https://www.anthropic.com/constitutionOpenAI Preparedness Framework v2, Table 5
“Monitor AI supervises agent actions.”
https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report, August 2026 / alignment-assessment brief / Claude's Constitution | “"we prompted an instance of Claude Mythos 5 to review a near-final draft of Section 2 of this report"; access to "internal Anthropic Slack channels", "internal documents", "internal codebase" and subagents; "Readers should weigh my position honestly, as I do: I am a Claude model reviewing Anthropic's assessment of Claude models"; "In practice, Claude took 24 minutes to produce this review" (§2.20). "a prompted Claude model reviews suggested code changes" (§2.23.2.3). Incident scan second stage "used Claude to review the '9.2 million transcripts'". Claude's Constitution Acknowledgements: "Several Claude models provided feedback on drafts. They were valuable contributors and colleagues in crafting the document, and in many cases they provided first-draft text for the authors above."” Also cites {ALA}, {CONST}. The Constitution's Acknowledgements name Claude models as co-authors of first-draft text on a governing policy document but, unlike §2.20, do not name the accountable human editor(s) of that text. | exact confidence: high | Anthropic Risk Report, August 2026 |
| OpenAI OpenAI Preparedness Framework v2 / GPT-5.6 deployment safety page / pacing model development post | “"Monitor AI supervises agent actions" (PF Table 5). "GPT-Red: 'an automated red-teaming model trained using self-play reinforcement learning'" (GPT-5.6 s.4.2). "Automated investigators" (August 2026).” Also cites {G56}, {OPACE}. | close confidence: medium | OpenAI Preparedness Framework v2 |
| Google DeepMind Gemini 3.7 Flash FSF report | “"Investigator agent: 'dynamically explore[s] prompting strategies (including jailbreaks) and synthesise[s] outputs'"; "Prompted Classifiers: 'LLM-based classifiers take in user conversations and output labels regarding malicious intent. Developed using AlphaEvolve'" (pp.27-28).” | close confidence: medium | Gemini 3.7 Flash FSF report |
| xAI Grok 4.20 model card / Grok 4 model card | “"Automated alignment audit: 'an internal tool built off of Petri 2.0 ... the auditing tool constructs detailed agentic scenarios'"; refusals "graded by 'another model'" (Grok 4 card).” Also cites {G4}. | close confidence: medium | Grok 4.20 model card |
| Meta Meta Advanced AI Scaling Framework v2 | “LlamaFirewall chain-of-thought auditing; risk of "long-term dependency of AI for supervisory functions" (§4.2.4).” | narrow confidence: medium | Meta Advanced AI Scaling Framework v2 |
| METR METR OpenAI-Hugging Face incident investigation / METR (site) | “Investigation relied on GPT-5.6 Sol "analysis agents"; METR asks: "Is there evidence that the AI tools used for scanning and filtering through the results from our queries might fail to surface serious misalignment incidents (whether due to benign failures or intentional sandbagging)?"” Also cites {METR}. | close confidence: medium | METR: OpenAI-Hugging Face incident investigation |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- ICMJE (discovery)
"Authors should not list or cite AI and AI-assisted technologies as an author" (discovery) -- RL, a pointer not a mapping; concerns human-authorship credit, adjacent to but out of scope for this evaluator-independence cluster.
Gap
*Accountability note (source document, revised in a later pass):* CRediT roles cover research outputs and have no representation for non-human contributors. Anthropic's §2.20 is the only instance in the corpus that records an AI reviewer's access, time on task, self-declared conflict and the disposition of its criticisms -- an operational, task-level record. Claude's Constitution is the closer analogue to a CRediT statement: it names an AI system class ("Several Claude models") as a contributor to a specific governing document, with a stated contribution type, the way a CRediT byline credits a contributor role -- but without this element's other fields (which model instance, what access, what human supervised or accepted the AI-drafted text). METR's own question shows the same AI tools can compromise an investigation (in the OpenAI case, agents "successfully spoofed tool calls in METR's own transcripts"). H.R. 9925 (FRONTIER Act, not enacted) supplies a third, statutory-drafting-stage analogue: its required compliance-audit report must include "a list of personnel" involved, which its own research brief calls "the closest the bill comes to contributor credit" (Sec. 4(c)) -- a named-personnel disclosure requirement, not a role-typed CRediT-style byline, and the bill nowhere contemplates an AI system as a contributor to the audit or assessment work itself.







