Skip to main content
v2026.11,858 entries · CC-BY 4.0
NIKOLAI elementN6 · Mitigations and securityProposednikolai-v0.1

Robustness level

NIKOLAI proposes Robustness level as a graded controlled-value axis stating how hard a safeguard is to circumvent, paired with the attack model assumed and the evidence supporting the grade (e.g. red-team success rate, independently verified multiplier). This is an unsourced NIKOLAI editorial synthesis; Anthropic and Google DeepMind both maintain a named, graded robustness/adversarial-robustness axis, which supports treating it as a distinct controlled-value ladder from coverage level.

This is CASRAI's own proposed definition, not a definition any named organisation has agreed to. See what NIKOLAI is and is not.

Source of record

Where this definition comes from

Crosswalk

How named organisations use this concept

Every row below is a shadow mapping. It is CASRAI's own reading of a published document. No lab, evaluator or regulator named here has declared, endorsed, or been consulted on this mapping. That will change only when an organisation files its own Mapping Declaration — see the non-endorsement policy.
OrganisationTheir term, as publishedMatchSource
Anthropic
Anthropic Risk Report, August 2026
Level 1 = the design from the previous Risk Report, patched for known jailbreaks; Level 2 = updated Constitutional Classifiers (Jan 2026 paper), streaming linear probe + fine-tuned second stage, weighted combination; Level 3 = Level 2 "with a lower threshold for blocking queries in order to be more conservative for our most capable models"exact
confidence: high
Anthropic Risk Report, August 2026
OpenAI
OpenAI Preparedness Framework v2 / GPT-5.6
"Robustness (claim): Users cannot use the model to cause the harm because they cannot elicit the capability, such as because the model is modified to refuse to provide assistance to harmful tasks and is robust to jailbreaks that would circumvent those refusals." Efficacy metrics include "Time to patching a new known jailbreak"; GPT-5.6 "Worst-case defender success rate"close
confidence: medium
OpenAI Preparedness Framework v2
Google DeepMind
Gemini 3.7 Flash FSF Report
"Adversarial Robustness: 'if threat actors do try to circumvent safeguards on "covered" topics, how well would they be able to do so?'"; "Violation Rate"close
confidence: medium
Gemini 3.7 Flash FSF Report
Meta
Meta Advanced AI Scaling Framework v2
CB acceptance: "40% refusal or safe responses against all adversarial attacks within a typical adversarial attack portfolio"narrow
confidence: medium
Meta Advanced AI Scaling Framework v2
G42
G42 Frontier Safety Framework
"DML 2" objective: "Even a determined actor should not be able to reliably elicit CBRN weapons advice ..."
Source document flags this row assignment as an open [VERIFY] item (ALIGNMENT-MATRIX.md §6 item 16, "G42 DML 2 row assignment" — not resolved as of the 16 Sep 2026 pass); treat the row placement itself, not just the content, as provisional.
close
confidence: low
G42 Frontier Safety Framework

Related, not mapped

Pointers that are not crosswalk claims

These sources mention this concept but do not define or map it clearly enough to count as a crosswalk row — noted here so the research is visible without overstating it as a mapping.

  • xAI

    HackerBench "under 'the standard release-tracked safeguards': 6.9% harmful/dual-use compliance and 0.0% benign refusal" — related benchmark result, not a mapping (RL).

    Grok 4.6 Model Card
  • UK AISI (via Anthropic)

    "Boundary Point Jailbreaking (BPJ)": "an automated methodology developed by UK AISI that optimizes against black-box classifiers"; AISI "independently verified the Level 3 robustness multiple (AISI estimated 2x, Anthropic 3x)" — related methodology/independent-verification note, not a mapping (RL).

    Anthropic Risk Report, August 2026
  • Frontier Model Forum

    "Adversarial testing: structured attempts to circumvent model safeguards and elicit harmful behaviors through techniques that malicious actors might employ." — related, not a mapping (RL).

    Frontier Model Forum, Third-Party Assessments

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →