Evidence and evaluations
Evaluation, evaluation run, elicitation method, saturation status, evaluation-validity threat.
Every capability threshold rests on a dangerous-capability evaluation, and an evaluation's evidentiary weight depends on how hard the model was pushed to perform and whether the benchmark still discriminates capability. This track defines evaluation, elicitation method, and saturation status -- the methodology vocabulary behind evaluation runs at Anthropic, OpenAI, Google DeepMind, METR, and the Frontier Model Forum, and behind the EU GPAI Code of Practice's own evaluation requirements.
- Evaluation-validity threatProposedcontrolled-value
A proposed controlled list of conditions, such as sandbagging or reward hacking, where a model's behavior in evaluation may differ from deployment.
- Saturation statusProposedcontrolled-value
A proposed value recording whether an evaluation can still discriminate capability, the criterion used, and whether to retire, replace, or keep it.
- Elicitation methodProposedproperty
A proposed record of the techniques used to draw out a model's maximum capability in an evaluation, plus whether the result is a lower bound or ceiling.
- Evaluation runProposedrecord-type
A proposed record of one execution, or batch of executions, of an evaluation against a specific model checkpoint, with its attempt count and scoring rule.
- EvaluationProposedrecord-type
A proposed definition of an evaluation: a defined procedure producing evidence about a model's capabilities, propensities, or safeguard effectiveness.







