Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN5 · Evidence and evaluationsProposednikolai-v0.1

Elicitation method

NIKOLAI proposal: the techniques and conditions used to draw out a model's maximum capability during an evaluation, recorded as a property of the evaluation run, together with the declared interpretation of the result (lower bound vs. ceiling).

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. It is CASRAI's own reading of a published document. No lab, evaluator or regulator named here has declared, endorsed, or been consulted on this mapping. That will change only when an organisation files its own Mapping Declaration — see the non-endorsement policy.
OrganisationTheir term, as publishedMatchSource
Anthropic
Anthropic Risk Report (August 2026)
Elicitation experiments raised Mythos 5 "from 0% to 3.8% (fine-tuning) and 9.2% (prompt optimisation)" (§2.7); "we have not provided clear evidence that this elicitation is sufficiently strong" (§2.16.1).close
confidence: high
Anthropic Risk Report (August 2026)
OpenAI
OpenAI Preparedness Framework v2
Elicitation aims at "the high end of expected elicitation by threat actors" using highest-capability settings, a variant with a "negligible rate of safety-based refusals", best scaffolds, and fine-tuning if weights will be released; "we regard any one-time capability elicitation in a frontier model as a lower bound, rather than a ceiling" (§3.1).exact
confidence: high
OpenAI Preparedness Framework v2
OpenAI
GPT-5.6 deployment safety card
SecureBio evaluated "a railfree version of GPT-5.6 Sol" with filters disabled (s.9.1.1.8).exact
confidence: high
GPT-5.6 deployment safety card
Google DeepMind
Gemini 3.7 Flash FSF report
"Elicitation: methods 'to surface a model's latent capabilities and propensities, and ensure that our risk assessment is based on the model's absolute potential rather than its default behavior'" (p.5).exact
confidence: high
Gemini 3.7 Flash FSF report
Google DeepMind
Frontier Safety Framework v3.1
evaluations use "appropriate scaffolding, inference compute, and other augmentations" (s.1.3.2).exact
confidence: high
Google DeepMind Frontier Safety Framework v3.1
xAI
Grok 4.6 model card
"Unrestricted configuration": capability testing "without the safeguards we use in production, which would otherwise mask the model's full capability" (§7, §7.1) — safeguard removal only, narrower than the general elicitation-method concept.narrow
confidence: medium
Grok 4.6 model card
Meta
Meta Advanced AI Scaling Framework v2
Helpful-only fine-tuning "(i.e., refusal-free)", domain-specific capability training for open release, task-optimised scaffolds; "Agents will be provided with a generous token budget, up to the maximum context length or to the point at which performance plateaus" (§4.2).exact
confidence: high
Meta Advanced AI Scaling Framework v2
METR
METR
"Full Capability Elicitation During Evaluations: Intentions to perform model evaluations in a way that does not underestimate the full capabilities of the model."exact
confidence: high
METR
EU
EU GPAI Code of Practice, Safety and Security Chapter
Appendix 3.2 "Model elicitation": evaluations must use "at least a state-of-the-art level of model elicitation that elicits the model's capabilities, propensities, affordances, and/or effects", techniques that "(1) minimise the risk of under-elicitation; and (2) minimise the risk of model deception during model evaluations (e.g. sandbagging)"; Appendix 3.4 requires access to "the model version(s) with the fewest safety mitigations implemented (such as a helpful-only model version, if it exists)" and an indicative time floor of "at least 20 business days ... for most systemic risks and model evaluation methods."exact
confidence: high
EU GPAI Code of Practice, Safety and Security Chapter

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • UK Government / AISI

    UK AISI "Grey box access": "access to chain-of-thought of the safety reasoning monitor, exact policy wording, and real-time feedback on classifier labels that would not be accessible to real-world attackers" — an access-tier pointer, not itself an elicitation-method definition.

    GPT-5.6 deployment safety card
  • Frontier Model Forum

    Access item "version with limited safety training" — a pointer, not a full elicitation-method definition.

    Frontier Model Forum, Third-Party Assessments

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →