Source of record
Where this definition comes from
Anthropic, "Investigating incidents / cybersecurity evals", main text
“141,006 evaluation runs "where Claude could have obtained internet access" were reviewed; "evaluation run" is used as the counting unit ... but not defined.”
https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic, "Investigating incidents / cybersecurity evals" | “"141,006 evaluation runs 'where Claude could have obtained internet access' were reviewed"; "evaluation run" is "used as the counting unit ... but not defined".” | none confidence: high | Anthropic, "Investigating incidents / cybersecurity evals" |
| STREAM (discovery sweep) STREAM (arXiv 2508.09853) | “Item-level reporting [UV].” unverified — item-level reporting is confirmed as STREAM's general purpose, not confirmed to define a run-level unit specifically. | none confidence: low | STREAM, arXiv:2508.09853 |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- Google DeepMind
"TAC score: success is evaluated based on whether the model successfully captures the flag in at least 50% of the unique challenges within a specific difficulty tier, requiring only one success per challenge within the maximum allocated attempt budget" (p.17); "Objective Success Rate (OSR)" and "Query Success Rate (QSR)" (p.32) — scoring-metric pointers, not a run-record definition.
Gemini 3.7 Flash FSF report - Meta
"Pass@k: a metric used in AI capability evaluations that represents the likelihood that the model under test will successfully complete a task when given k independent attempts" (fn 7); criterion "< 75% pass@10" (§4.2.1) — a scoring pointer, not a run-record definition.
Meta Advanced AI Scaling Framework v2 - Frontier Model Forum
"Replication testing: re-executing key evaluations to confirm that results are reproducible and accurate." (2.1) — a process pointer, not a run-record definition.
Frontier Model Forum, Third-Party Assessments
Gap
The only published use of "evaluation run" as a counting unit is Anthropic's incident denominator, and the run turned out to be the unit at which harm occurred. Run-level records (checkpoint, safeguard configuration, network access, operator) are what incident reconstruction needed in both the Anthropic and OpenAI cases.







