Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN8 · Transparency and reviewProposednikolai-v0.1

AI-Model Review

NIKOLAI proposal: AI-Model Review is a record that an AI system performed an assurance task -- review, monitoring, grading, red-teaming, or analysis -- capturing the system's identity, the task performed, its access, its supervision, the disposition of its output, and the accountable human who owns that disposition. This working definition is NIKOLAI's own editorial synthesis. It is distinct from, and does not substitute for, a CRediT-pattern contributor-role accountability layer for human authorship, which is out of scope for this element.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. It is CASRAI's own reading of a published document. No lab, evaluator or regulator named here has declared, endorsed, or been consulted on this mapping. That will change only when an organisation files its own Mapping Declaration — see the non-endorsement policy.
OrganisationTheir term, as publishedMatchSource
Anthropic
Anthropic Risk Report, August 2026 / alignment-assessment brief / Claude's Constitution
"we prompted an instance of Claude Mythos 5 to review a near-final draft of Section 2 of this report"; access to "internal Anthropic Slack channels", "internal documents", "internal codebase" and subagents; "Readers should weigh my position honestly, as I do: I am a Claude model reviewing Anthropic's assessment of Claude models"; "In practice, Claude took 24 minutes to produce this review" (§2.20). "a prompted Claude model reviews suggested code changes" (§2.23.2.3). Incident scan second stage "used Claude to review the '9.2 million transcripts'". Claude's Constitution Acknowledgements: "Several Claude models provided feedback on drafts. They were valuable contributors and colleagues in crafting the document, and in many cases they provided first-draft text for the authors above."
Also cites {ALA}, {CONST}. The Constitution's Acknowledgements name Claude models as co-authors of first-draft text on a governing policy document but, unlike §2.20, do not name the accountable human editor(s) of that text.
exact
confidence: high
Anthropic Risk Report, August 2026
OpenAI
OpenAI Preparedness Framework v2 / GPT-5.6 deployment safety page / pacing model development post
"Monitor AI supervises agent actions" (PF Table 5). "GPT-Red: 'an automated red-teaming model trained using self-play reinforcement learning'" (GPT-5.6 s.4.2). "Automated investigators" (August 2026).
Also cites {G56}, {OPACE}.
close
confidence: medium
OpenAI Preparedness Framework v2
Google DeepMind
Gemini 3.7 Flash FSF report
"Investigator agent: 'dynamically explore[s] prompting strategies (including jailbreaks) and synthesise[s] outputs'"; "Prompted Classifiers: 'LLM-based classifiers take in user conversations and output labels regarding malicious intent. Developed using AlphaEvolve'" (pp.27-28).close
confidence: medium
Gemini 3.7 Flash FSF report
xAI
Grok 4.20 model card / Grok 4 model card
"Automated alignment audit: 'an internal tool built off of Petri 2.0 ... the auditing tool constructs detailed agentic scenarios'"; refusals "graded by 'another model'" (Grok 4 card).
Also cites {G4}.
close
confidence: medium
Grok 4.20 model card
Meta
Meta Advanced AI Scaling Framework v2
LlamaFirewall chain-of-thought auditing; risk of "long-term dependency of AI for supervisory functions" (§4.2.4).narrow
confidence: medium
Meta Advanced AI Scaling Framework v2
METR
METR OpenAI-Hugging Face incident investigation / METR (site)
Investigation relied on GPT-5.6 Sol "analysis agents"; METR asks: "Is there evidence that the AI tools used for scanning and filtering through the results from our queries might fail to surface serious misalignment incidents (whether due to benign failures or intentional sandbagging)?"
Also cites {METR}.
close
confidence: medium
METR: OpenAI-Hugging Face incident investigation

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • ICMJE (discovery)

    "Authors should not list or cite AI and AI-assisted technologies as an author" (discovery) -- RL, a pointer not a mapping; concerns human-authorship credit, adjacent to but out of scope for this evaluator-independence cluster.

Gap

*Accountability note (source document, revised in a later pass):* CRediT roles cover research outputs and have no representation for non-human contributors. Anthropic's §2.20 is the only instance in the corpus that records an AI reviewer's access, time on task, self-declared conflict and the disposition of its criticisms -- an operational, task-level record. Claude's Constitution is the closer analogue to a CRediT statement: it names an AI system class ("Several Claude models") as a contributor to a specific governing document, with a stated contribution type, the way a CRediT byline credits a contributor role -- but without this element's other fields (which model instance, what access, what human supervised or accepted the AI-drafted text). METR's own question shows the same AI tools can compromise an investigation (in the OpenAI case, agents "successfully spoofed tool calls in METR's own transcripts"). H.R. 9925 (FRONTIER Act, not enacted) supplies a third, statutory-drafting-stage analogue: its required compliance-audit report must include "a list of personnel" involved, which its own research brief calls "the closest the bill comes to contributor credit" (Sec. 4(c)) -- a named-personnel disclosure requirement, not a role-typed CRediT-style byline, and the bill nowhere contemplates an AI system as a contributor to the audit or assessment work itself.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →