Source of record
Where this definition comes from
OpenAI Preparedness Framework v2, §3.1
“we regard any one-time capability elicitation in a frontier model as a lower bound, rather than a ceiling.”
https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdfEU GPAI Code of Practice, Safety and Security Chapter, Appendix 3.2, 3.4
“at least a state-of-the-art level of model elicitation that elicits the model's capabilities, propensities, affordances, and/or effects ... (1) minimise the risk of under-elicitation; and (2) minimise the risk of model deception during model evaluations (e.g. sandbagging).”
https://ec.europa.eu/newsroom/dae/redirection/document/118119Gemini 3.7 Flash FSF report, p.5
“Elicitation: methods "to surface a model's latent capabilities and propensities, and ensure that our risk assessment is based on the model's absolute potential rather than its default behavior".”
https://storage.googleapis.com/deepmind-media/gemini/gemini_3-7_flash_fsf_report.pdf
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report (August 2026) | “Elicitation experiments raised Mythos 5 "from 0% to 3.8% (fine-tuning) and 9.2% (prompt optimisation)" (§2.7); "we have not provided clear evidence that this elicitation is sufficiently strong" (§2.16.1).” | close confidence: high | Anthropic Risk Report (August 2026) |
| OpenAI OpenAI Preparedness Framework v2 | “Elicitation aims at "the high end of expected elicitation by threat actors" using highest-capability settings, a variant with a "negligible rate of safety-based refusals", best scaffolds, and fine-tuning if weights will be released; "we regard any one-time capability elicitation in a frontier model as a lower bound, rather than a ceiling" (§3.1).” | exact confidence: high | OpenAI Preparedness Framework v2 |
| OpenAI GPT-5.6 deployment safety card | “SecureBio evaluated "a railfree version of GPT-5.6 Sol" with filters disabled (s.9.1.1.8).” | exact confidence: high | GPT-5.6 deployment safety card |
| Google DeepMind Gemini 3.7 Flash FSF report | “"Elicitation: methods 'to surface a model's latent capabilities and propensities, and ensure that our risk assessment is based on the model's absolute potential rather than its default behavior'" (p.5).” | exact confidence: high | Gemini 3.7 Flash FSF report |
| Google DeepMind Frontier Safety Framework v3.1 | “evaluations use "appropriate scaffolding, inference compute, and other augmentations" (s.1.3.2).” | exact confidence: high | Google DeepMind Frontier Safety Framework v3.1 |
| xAI Grok 4.6 model card | “"Unrestricted configuration": capability testing "without the safeguards we use in production, which would otherwise mask the model's full capability" (§7, §7.1) — safeguard removal only, narrower than the general elicitation-method concept.” | narrow confidence: medium | Grok 4.6 model card |
| Meta Meta Advanced AI Scaling Framework v2 | “Helpful-only fine-tuning "(i.e., refusal-free)", domain-specific capability training for open release, task-optimised scaffolds; "Agents will be provided with a generous token budget, up to the maximum context length or to the point at which performance plateaus" (§4.2).” | exact confidence: high | Meta Advanced AI Scaling Framework v2 |
| METR METR | “"Full Capability Elicitation During Evaluations: Intentions to perform model evaluations in a way that does not underestimate the full capabilities of the model."” | exact confidence: high | METR |
| EU EU GPAI Code of Practice, Safety and Security Chapter | “Appendix 3.2 "Model elicitation": evaluations must use "at least a state-of-the-art level of model elicitation that elicits the model's capabilities, propensities, affordances, and/or effects", techniques that "(1) minimise the risk of under-elicitation; and (2) minimise the risk of model deception during model evaluations (e.g. sandbagging)"; Appendix 3.4 requires access to "the model version(s) with the fewest safety mitigations implemented (such as a helpful-only model version, if it exists)" and an indicative time floor of "at least 20 business days ... for most systemic risks and model evaluation methods."” | exact confidence: high | EU GPAI Code of Practice, Safety and Security Chapter |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- UK Government / AISI
UK AISI "Grey box access": "access to chain-of-thought of the safety reasoning monitor, exact policy wording, and real-time feedback on classifier labels that would not be accessible to real-world attackers" — an access-tier pointer, not itself an elicitation-method definition.
GPT-5.6 deployment safety card - Frontier Model Forum
Access item "version with limited safety training" — a pointer, not a full elicitation-method definition.
Frontier Model Forum, Third-Party Assessments







