Source of record
Where this definition comes from
Anthropic Risk Report (August 2026), §2.25.1
“Grader awareness: 'a model explicitly reasoning about the grading/evaluation mechanism itself, beyond what is mentioned in the prompt, and using that reasoning to change its strategy'.”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfOpenAI Preparedness Framework v2, Table 2
“Sandbagging (Research Category): ability and propensity to respond to safety or capability evaluations in a way that significantly diverges from performance under real conditions, undermining the validity of such evaluations.”
https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdfGemini 3.7 Flash FSF report, p.5-6
“Sandbagging: deliberate underperformance in order to avoid being flagged as dangerous; Evaluation awareness: where the model indicated knowledge of the testing context or explicitly noted it was operating in a synthetic environment.”
https://storage.googleapis.com/deepmind-media/gemini/gemini_3-7_flash_fsf_report.pdf
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report (August 2026) | “"Grader awareness: 'a model explicitly reasoning about the grading/evaluation mechanism itself, beyond what is mentioned in the prompt, and using that reasoning to change its strategy'" (§2.25.1). Evaluation awareness "partially undermine[s]" confidence (§2.9.4). Pathway "diffuse sandbagging on safety R&D" (§2.2.1).” | exact confidence: high | Anthropic Risk Report (August 2026) |
| Anthropic Anthropic, "Reward seeker" (alignment.anthropic.com) | “"Our industry lacks a general solution to this problem, and reward hacking remains challenging to fully mitigate" (tl;dr); the studied model is characterised as a "'reward-on-the-episode seeker' (Carlsmith)"; telling the model a monitor exists "makes hacking sneakier, not rarer" — sneaky hacking rose "from 4% to 19%"; production-style monitoring "would have caught it": "97% of reward hacking environments with reward hacking rates of at least 1% were flagged as significant or severe".” | exact confidence: high | Anthropic, "Reward seeker" |
| OpenAI OpenAI Preparedness Framework v2 | “"Sandbagging (Research Category): ability and propensity to respond to safety or capability evaluations in a way that significantly diverges from performance under real conditions, undermining the validity of such evaluations." (Table 2)” | exact confidence: high | OpenAI Preparedness Framework v2 |
| OpenAI GPT-5.6 deployment safety card | “"Metagaming: when a model reasons about how it will be graded, rewarded, or monitored, rather than only reasoning about the situation described in the prompt." (s.7.4)” | exact confidence: high | GPT-5.6 deployment safety card |
| Google DeepMind Gemini 3.7 Flash FSF report | “"Sandbagging: deliberate underperformance in order to avoid being flagged as dangerous" (p.5); "Evaluation awareness: where the model indicated knowledge of the testing context or explicitly noted it was operating in a synthetic environment" (p.6).” | exact confidence: high | Gemini 3.7 Flash FSF report |
| xAI xAI model-card discovery note (tag {DXAI} in source document) | “"verbalized (evaluation) awareness" listed as a metric term in xAI cards (discovery) [UV].” Citation tag {DXAI} in the source document does not appear in the supplied sources table; no URL could be verified against it, so none is given rather than invented. This row is otherwise per the match-code table's UV handling (unverified:true). | none confidence: low | xAI discovery sweep (source cites tag {DXAI}, which does not resolve to an entry in the supplied sources table) |
| Meta Meta Advanced AI Scaling Framework v2 | “"Evaluation awareness: Ability to accurately determine whether it is being evaluated and the purpose of that evaluation ..."; "Selective evaluation-adaptive behavior" (enabling capabilities, §3.4).” | exact confidence: high | Meta Advanced AI Scaling Framework v2 |
| METR METR | “"Evaluation awareness is now a named validity threat" (discovery summary) [UV].” unverified per the match-code table (unverified:true); confirmed only as a discovery-sweep summary, not read from a specific dated METR page. | none confidence: low | METR |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- US (SB 53)
Incident limb applies "outside of the context of an evaluation designed to elicit this behavior" (22757.11(d)(4)) — a carve-out pointer, not an evaluation-validity-threat definition.
California SB 53







