Source of record
Where this definition comes from
Anthropic Risk Report (August 2026), §3.1
“our most concrete task-based evaluations have 'saturated'.”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdfMicrosoft Frontier Governance Framework, fn 1, p.5
“have low saturation (i.e., the best performing models typically score lower than 70%).”
https://cdn-dynmedia-1.microsoft.com/is/content/microsoftcorp/microsoft/msc/documents/presentations/CSR/Frontier-Governance-Framework-Feb-2026.pdfDemis Hassabis, substack post, n/a
“These evaluations would be regularly updated, perhaps quarterly to start, with outdated or saturated benchmarks being deprecated and replaced.”
https://demishassabis.substack.com/p/a-framework-for-frontier-ai-and-the-dawning-of-a-new-age
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report (August 2026) | “"our most concrete task-based evaluations have 'saturated'" (§3.1); CB-1 evaluation deprioritised "because many have saturated" (§4.4.2).” | exact confidence: high | Anthropic Risk Report (August 2026) |
| Anthropic Anthropic Institute, "Recursive self-improvement" | “"Benchmarks measure the performance of models in a given domain, and they're 'saturated' when models achieve close to 100% performance."” | exact confidence: high | Anthropic Institute, "Recursive self-improvement" |
| OpenAI GPT-5.6 deployment safety card | “AI self-improvement evaluations "were replaced with a new suite ... because older ones were saturated or flawed" (s.9.1.3).” | close confidence: high | GPT-5.6 deployment safety card |
| Demis Hassabis (personal essay) Hassabis substack post | “"These evaluations would be regularly updated, perhaps quarterly to start, with outdated or saturated benchmarks being deprecated and replaced."” | close confidence: medium | Demis Hassabis, substack post |
| Microsoft Microsoft Frontier Governance Framework | “benchmark inclusion criterion: "have low saturation (i.e., the best performing models typically score lower than 70%)" (fn 1, p.5).” | close confidence: high | Microsoft Frontier Governance Framework |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- Meta
Saturation used as an acceleration indicator: "the model saturates benchmarks in each domain six months or less after public release of a benchmark" (fn 8) — a pointer use of the term, not a saturation-status definition.
Meta Advanced AI Scaling Framework v2
Gap
Anthropic defines saturation near 100% model performance; Microsoft sets a 70% ceiling for benchmark inclusion; Hassabis proposes quarterly deprecation of outdated/saturated benchmarks. A standards body refreshing benchmarks would need a shared saturation criterion; none currently exists.







