Source of record
Where this definition comes from
OpenAI Preparedness Framework v2, §3.1
“Scalable Evaluations: automated evaluations designed to measure proxies that approximate whether a capability threshold has been crossed.”
https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdfGoogle DeepMind Frontier Safety Framework v3.1, Glossary
“Early Warning Evaluations: are evaluations which measure the dangerous capabilities of a model. They specifically target the threats and risk scenarios identified through our threat modeling to determine a model's proximity to a CCL or TCL.”
https://storage.googleapis.com/deepmind-media/DeepMind.com/Blog/strengthening-our-frontier-safety-framework/frontier-safety-framework_3-1.pdfEU GPAI Code of Practice, Safety and Security Chapter, Measure 3.2
“Signatories will conduct at least state-of-the-art model evaluations in the modalities relevant to the systemic risk to assess the model's capabilities, propensities, affordances, and/or effects.”
https://ec.europa.eu/newsroom/dae/redirection/document/118119
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report (August 2026) | “CoBench: "an internal evaluation measuring how well a model, placed at a historical point in Anthropic's infrastructure ..., can diagnose the root causes of issues that Anthropic engineers actually solved" (§3.4.3).” | exact confidence: high | Anthropic Risk Report (August 2026) |
| Anthropic Anthropic Advanced AI Framework | “the safety framework must describe "capability evaluations performed" (p.5).” | exact confidence: high | Anthropic Advanced AI Framework |
| OpenAI OpenAI Preparedness Framework v2 | “"Scalable Evaluations: automated evaluations designed to measure proxies that approximate whether a capability threshold has been crossed." "Deep Dives: designed to provide additional evidence validating the scalable evaluations' findings ..." (§3.1)” | exact confidence: high | OpenAI Preparedness Framework v2 |
| Google DeepMind Frontier Safety Framework v3.1 | “"Early Warning Evaluations: are evaluations which measure the dangerous capabilities of a model. They specifically target the threats and risk scenarios identified through our threat modeling to determine a model's proximity to a CCL or TCL" (glossary).” | exact confidence: high | Google DeepMind Frontier Safety Framework v3.1 |
| xAI xAI Frontier AI Framework (30 Jun 2026) | “"Model evaluations: state-of-the-art model evaluations relevant to the systemic risk to assess the model's capabilities, propensities, affordances, and/or effects, which may include Q&A sets, task-based evaluations, benchmarks, red-teaming and other methods of adversarial testing, human uplift studies, model organisms, simulations, and/or proxy evaluations for classified materials" (s.2.2(2)).” CORRECTION: this row cites {FAIF26} but the original extraction omitted the required draft-document caveat. Adding it now: this document's PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement disambiguating draft vs. final status was found (also unresolved per ALIGNMENT-MATRIX.md §6 item 4). Treat as provisional. | exact confidence: medium | xAI Frontier AI Framework (30 Jun 2026) |
| xAI Grok 4.6 model card | “"safety-threshold evaluations"” | exact confidence: high | Grok 4.6 model card |
| Meta Meta Advanced AI Scaling Framework v2 | “"Evaluation(s): refers to the assessments we do to understand capabilities and performance. We use this term to describe automated and human evaluations that assess capabilities, as well as evaluations to assess potential for misuse, such as red teaming and uplift studies." (Appendix I)” | exact confidence: high | Meta Advanced AI Scaling Framework v2 |
| EU EU GPAI Code of Practice, Safety and Security Chapter | “Measure 3.2: "Signatories will conduct at least state-of-the-art model evaluations in the modalities relevant to the systemic risk to assess the model's capabilities, propensities, affordances, and/or effects", including "open-ended testing ... with a view to identifying unexpected behaviours, capability boundaries, or emergent properties"; Appendix 3.1 requires "internal validity", "external validity" and "reproducibility".” | exact confidence: high | EU GPAI Code of Practice, Safety and Security Chapter |
| EU / xAI (cross-reference) EU CoP Appendix 3 example methods, echoed in xAI FAIF s.2.2(2) | “example methods (Appendix 3) "echoed by xAI s.2.2(2) almost verbatim": Q&A sets, task-based evaluations, benchmarks, red-teaming, human uplift studies, model organisms, simulations, proxy evaluations for classified materials.” This document's PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement was found disambiguating draft vs. final status as of this pass. Treat as provisional. | exact confidence: medium | xAI Frontier AI Framework (30 Jun 2026) |
| California SB 53 California SB 53 | “Large developers summarise "(A) Assessments of catastrophic risks from the frontier model conducted pursuant to the large frontier developer's frontier AI framework" (22757.12(c)(2)).” | broad confidence: medium | California SB 53 |
| US Government (NIST CAISI) NIST CAISI bulletin | “CAISI "will conduct pre-deployment evaluations and targeted research" (para. 1) — undefined.” | none confidence: medium | NIST CAISI bulletin |
| Frontier Model Forum FMF Third-Party Assessments technical report | “"Capability Assessments: evaluate whether a model crosses any enabling capability thresholds or outcomes-based thresholds" (1.2).” | close confidence: medium | Frontier Model Forum, Third-Party Assessments |
| Safety Framework Cards (discovery) Safety Framework Cards (SSRN 7061798) | “"evaluation methodology" dimension [UV] — abstract-only, full text paywalled/unread.” unverified: full paper is SSRN account-gated; per the match-code table UV rows carry unverified:true. Open per ALIGNMENT-MATRIX.md §6 item 5. | none confidence: low | Safety Framework Cards (SSRN 7061798, abstract only) |
| STREAM (discovery sweep) STREAM (arXiv 2508.09853) | “"the only published item-level reporting standard for evaluations", with a template and gold-standard examples (discovery description) [UV].” unverified: only the abstract/scope has been confirmed from the primary source; the template's exact field names have not been read (ALIGNMENT-MATRIX.md §6 item 6). | none confidence: low | STREAM, arXiv:2508.09853 |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- METR
"Full Capability Elicitation During Evaluations"; "Timing and Frequency of Evaluations" — named as process elements, not an evaluation definition (RL, not a mapping).
METR







