Source of record
Where this definition comes from
Anthropic Risk Report, August 2026, §4.5.1
“Level 1 = the design from the previous Risk Report, patched for known jailbreaks; Level 2 = updated Constitutional Classifiers (Jan 2026 paper), streaming linear probe + fine-tuned second stage, weighted combination; Level 3 = Level 2 "with a lower threshold for blocking queries in order to be more conservative for our most capable models"”
https://www-cdn.anthropic.com/f61d49fa5596956a5dec75fea0e973bf6a6a8378/Redacted%20Risk%20Report%20August%202026%20.pdf
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Anthropic Risk Report, August 2026 | “Level 1 = the design from the previous Risk Report, patched for known jailbreaks; Level 2 = updated Constitutional Classifiers (Jan 2026 paper), streaming linear probe + fine-tuned second stage, weighted combination; Level 3 = Level 2 "with a lower threshold for blocking queries in order to be more conservative for our most capable models"” | exact confidence: high | Anthropic Risk Report, August 2026 |
| OpenAI OpenAI Preparedness Framework v2 / GPT-5.6 | “"Robustness (claim): Users cannot use the model to cause the harm because they cannot elicit the capability, such as because the model is modified to refuse to provide assistance to harmful tasks and is robust to jailbreaks that would circumvent those refusals." Efficacy metrics include "Time to patching a new known jailbreak"; GPT-5.6 "Worst-case defender success rate"” | close confidence: medium | OpenAI Preparedness Framework v2 |
| Google DeepMind Gemini 3.7 Flash FSF Report | “"Adversarial Robustness: 'if threat actors do try to circumvent safeguards on "covered" topics, how well would they be able to do so?'"; "Violation Rate"” | close confidence: medium | Gemini 3.7 Flash FSF Report |
| Meta Meta Advanced AI Scaling Framework v2 | “CB acceptance: "40% refusal or safe responses against all adversarial attacks within a typical adversarial attack portfolio"” | narrow confidence: medium | Meta Advanced AI Scaling Framework v2 |
| G42 G42 Frontier Safety Framework | “"DML 2" objective: "Even a determined actor should not be able to reliably elicit CBRN weapons advice ..."” Source document flags this row assignment as an open [VERIFY] item (ALIGNMENT-MATRIX.md §6 item 16, "G42 DML 2 row assignment" — not resolved as of the 16 Sep 2026 pass); treat the row placement itself, not just the content, as provisional. | close confidence: low | G42 Frontier Safety Framework |
Related, not mapped
Pointers that are not crosswalk claims
These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.
- xAI
HackerBench "under 'the standard release-tracked safeguards': 6.9% harmful/dual-use compliance and 0.0% benign refusal" — related benchmark result, not a mapping (RL).
Grok 4.6 Model Card - UK AISI (via Anthropic)
"Boundary Point Jailbreaking (BPJ)": "an automated methodology developed by UK AISI that optimizes against black-box classifiers"; AISI "independently verified the Level 3 robustness multiple (AISI estimated 2x, Anthropic 3x)" — related methodology/independent-verification note, not a mapping (RL).
Anthropic Risk Report, August 2026 - Frontier Model Forum
"Adversarial testing: structured attempts to circumvent model safeguards and elicit harmful behaviors through techniques that malicious actors might employ." — related, not a mapping (RL).
Frontier Model Forum, Third-Party Assessments







