Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN5 · Evidence and evaluationsProposednikolai-v0.1

Evaluation

NIKOLAI proposal: a defined procedure that produces evidence about a model's capabilities, propensities, or safeguard effectiveness, recorded with its purpose, method type, and reporting format as a distinct instance — independent of any single lab's own vocabulary for it.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. It is CASRAI's own reading of a published document. No lab, evaluator or regulator named here has declared, endorsed, or been consulted on this mapping. That will change only when an organisation files its own Mapping Declaration — see the non-endorsement policy.
OrganisationTheir term, as publishedMatchSource
Anthropic
Anthropic Risk Report (August 2026)
CoBench: "an internal evaluation measuring how well a model, placed at a historical point in Anthropic's infrastructure ..., can diagnose the root causes of issues that Anthropic engineers actually solved" (§3.4.3).exact
confidence: high
Anthropic Risk Report (August 2026)
Anthropic
Anthropic Advanced AI Framework
the safety framework must describe "capability evaluations performed" (p.5).exact
confidence: high
Anthropic Advanced AI Framework
OpenAI
OpenAI Preparedness Framework v2
"Scalable Evaluations: automated evaluations designed to measure proxies that approximate whether a capability threshold has been crossed." "Deep Dives: designed to provide additional evidence validating the scalable evaluations' findings ..." (§3.1)exact
confidence: high
OpenAI Preparedness Framework v2
Google DeepMind
Frontier Safety Framework v3.1
"Early Warning Evaluations: are evaluations which measure the dangerous capabilities of a model. They specifically target the threats and risk scenarios identified through our threat modeling to determine a model's proximity to a CCL or TCL" (glossary).exact
confidence: high
Google DeepMind Frontier Safety Framework v3.1
xAI
xAI Frontier AI Framework (30 Jun 2026)
"Model evaluations: state-of-the-art model evaluations relevant to the systemic risk to assess the model's capabilities, propensities, affordances, and/or effects, which may include Q&A sets, task-based evaluations, benchmarks, red-teaming and other methods of adversarial testing, human uplift studies, model organisms, simulations, and/or proxy evaluations for classified materials" (s.2.2(2)).
CORRECTION: this row cites {FAIF26} but the original extraction omitted the required draft-document caveat. Adding it now: this document's PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement disambiguating draft vs. final status was found (also unresolved per ALIGNMENT-MATRIX.md §6 item 4). Treat as provisional.
exact
confidence: medium
xAI Frontier AI Framework (30 Jun 2026)
xAI
Grok 4.6 model card
"safety-threshold evaluations"exact
confidence: high
Grok 4.6 model card
Meta
Meta Advanced AI Scaling Framework v2
"Evaluation(s): refers to the assessments we do to understand capabilities and performance. We use this term to describe automated and human evaluations that assess capabilities, as well as evaluations to assess potential for misuse, such as red teaming and uplift studies." (Appendix I)exact
confidence: high
Meta Advanced AI Scaling Framework v2
EU
EU GPAI Code of Practice, Safety and Security Chapter
Measure 3.2: "Signatories will conduct at least state-of-the-art model evaluations in the modalities relevant to the systemic risk to assess the model's capabilities, propensities, affordances, and/or effects", including "open-ended testing ... with a view to identifying unexpected behaviours, capability boundaries, or emergent properties"; Appendix 3.1 requires "internal validity", "external validity" and "reproducibility".exact
confidence: high
EU GPAI Code of Practice, Safety and Security Chapter
EU / xAI (cross-reference)
EU CoP Appendix 3 example methods, echoed in xAI FAIF s.2.2(2)
example methods (Appendix 3) "echoed by xAI s.2.2(2) almost verbatim": Q&A sets, task-based evaluations, benchmarks, red-teaming, human uplift studies, model organisms, simulations, proxy evaluations for classified materials.
This document's PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement was found disambiguating draft vs. final status as of this pass. Treat as provisional.
exact
confidence: medium
xAI Frontier AI Framework (30 Jun 2026)
California SB 53
California SB 53
Large developers summarise "(A) Assessments of catastrophic risks from the frontier model conducted pursuant to the large frontier developer's frontier AI framework" (22757.12(c)(2)).broad
confidence: medium
California SB 53
US Government (NIST CAISI)
NIST CAISI bulletin
CAISI "will conduct pre-deployment evaluations and targeted research" (para. 1) — undefined.none
confidence: medium
NIST CAISI bulletin
Frontier Model Forum
FMF Third-Party Assessments technical report
"Capability Assessments: evaluate whether a model crosses any enabling capability thresholds or outcomes-based thresholds" (1.2).close
confidence: medium
Frontier Model Forum, Third-Party Assessments
Safety Framework Cards (discovery)
Safety Framework Cards (SSRN 7061798)
"evaluation methodology" dimension [UV] — abstract-only, full text paywalled/unread.
unverified: full paper is SSRN account-gated; per the match-code table UV rows carry unverified:true. Open per ALIGNMENT-MATRIX.md §6 item 5.
none
confidence: low
Safety Framework Cards (SSRN 7061798, abstract only)
STREAM (discovery sweep)
STREAM (arXiv 2508.09853)
"the only published item-level reporting standard for evaluations", with a template and gold-standard examples (discovery description) [UV].
unverified: only the abstract/scope has been confirmed from the primary source; the template's exact field names have not been read (ALIGNMENT-MATRIX.md §6 item 6).
none
confidence: low
STREAM, arXiv:2508.09853

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • METR

    "Full Capability Elicitation During Evaluations"; "Timing and Frequency of Evaluations" — named as process elements, not an evaluation definition (RL, not a mapping).

    METR

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →