Skip to main content
v2026.11,858 entries · CC-BY 4.0

Security Level: Nine Labs and Regulators, One Undefined Standard

Google DeepMind and Anthropic both cite RAND’s security-level scale but define different tiers; xAI, Meta, and the EU’s AI Act code reject grading and require a stated goal instead; and OpenAI’s own framework names a ‘Critical’ standard it says it hasn’t specified yet.

Written and maintained by CASRAI Editorial Board

Last updated

Every frontier AI developer’s safety framework says, in one form or another, that model weights and training infrastructure need to be secured against theft. Almost none of them say it the same way. Google DeepMind and Anthropic both point to the same external benchmark — a RAND Corporation report on securing AI model weights — but define different tiers against it. OpenAI names a security standard tied to its highest capability tier and, by its own admission, hasn’t specified what that standard actually requires yet. xAI, Meta, and the EU’s AI Act code of practice reject graded scales entirely, requiring developers to state a goal instead: which attackers, specifically, are you defending against? G42 built its own four-level scale from scratch. Microsoft and METR each take a third approach again. Nine organizations, one underlying question — how secure is secure enough — and no shared answer. This is the first standalone deep-dive CASRAI’s NIKOLAI project has published on a single N6 element, and Security Level is a useful one to start with: it’s the element where the “graded scale vs. named goal” split that runs through the whole Mitigations and Security track is easiest to see clearly.

The RAND anchor: Google DeepMind’s SL2+, SL3, and SL4

Google DeepMind’s Frontier Safety Framework, version 3.1 (April 2026), is the most explicit of the nine about where its security tiers come from. Its own text states plainly: “We use security levels that indicate goals/principles in line with the corresponding level in the RAND framework” — a direct reference to RAND’s 2024 report Securing AI Model Weights. DeepMind then maps its own capability thresholds onto that RAND scale, level by level: CBRN uplift, cyber uplift, and harmful-manipulation capability thresholds each map to Security Level 2; an ML R&D acceleration threshold maps to Security Level 3; and for its highest capability threshold — full automation of Google’s own AI research teams — DeepMind writes that it “recommend[s] Security Level 4 for this capability threshold, but emphasize[s] that this must be taken on by the frontier AI field as a whole,” not by DeepMind alone.

DeepMind also defines a level RAND’s original scale doesn’t have a name for: “Security Level 2+,” which “uses RAND Security Level 2 (SL2) as a baseline, with additional security measures designed to address risks from insider threats and well-resourced non-state external actors” — things like dedicated insider-risk teams and background checks layered on top of RAND’s baseline external-attacker protections. Of the nine organizations covered here, DeepMind is the only one whose own document is this explicit about naming a specific RAND level for each capability tier.

Anthropic’s ASL-3: who’s covered, and who explicitly isn’t

Anthropic’s Responsible Scaling Policy (RSP), version 3.4, ties security requirements to its AI Safety Level (ASL) system rather than naming RAND levels directly for its own commitments — RAND only enters Anthropic’s document as an industry-wide recommendation, not Anthropic’s own company commitment. For its highest defined tier, aimed at threat actors “not bound by a credible governance regime,” the RSP states this “would likely mean security roughly in line with RAND SL4, but it depends on the capabilities of the strongest and most plausible threat actors.”

For Anthropic’s own ASL-3 Security Standard, the RSP’s change log is unusually direct about scope. The May 2025 update (RSP v2.2) states: “This update excludes both sophisticated insiders and state-compromised insiders from the ASL-3 Security Standard. Previously, only ‘highly sophisticated state-compromised insiders’ were explicitly excluded. […] the relatively small number of employees who might be capable of model theft does not significantly affect the risk level.” In practice, that means ASL-3 is built to hold against a well-resourced external attacker and an ordinary insider with legitimate access, but explicitly not against a sophisticated insider or a nation-state-backed one — the RSP itself states the underlying threat models for ASL-3’s current capability thresholds “do not warrant protection against either group.” The document elsewhere is direct about even the less extreme case: “Even malicious employees and other insiders with maximal levels of access will not be significantly enabled to cause catastrophic harm” is the outcome ASL-3 is built to guarantee, not an unlimited one.

OpenAI: a “Critical standard” the framework itself says isn’t specified yet

OpenAI’s Preparedness Framework ties security requirements to the same High/Critical capability-threshold language it uses throughout the document. For models crossing a High capability threshold, the framework requires (page 20) “robust security practices and controls,” listing named requirements — security threat modeling and risk management, defense in depth, access management, secure development and supply chain, operational security, and auditing and transparency — as the concrete content of that standard.

For the Critical tier, the framework is candid that the equivalent standard doesn’t exist yet. Table 1’s safeguard guidance for Critical capability thresholds states: “Until we have specified safeguards and security controls that would meet a Critical standard, halt further development.” That’s not a hedge from a secondary summary — it’s the Preparedness Framework’s own safeguard guideline, in its own words: a named “Critical standard” is invoked as the bar a model would need to clear, and the framework commits to pausing development rather than deploying without one, precisely because that standard hasn’t been written down. Of the nine approaches compared here, this is the clearest case of an organization naming a term for its highest security tier without yet defining it.

The goal-based camp: xAI, Meta, and the EU’s GPAI Code reject the graded scale

Three of the nine approaches skip graded levels entirely and require a stated goal instead.

The EU AI Act’s General-Purpose AI Code of Practice is the most explicit about this choice. Its Measure 6.1, titled “Security Goal,” states: “Signatories will define a goal that specifies the threat actors that their security mitigations are intended to protect against (‘Security Goal’), including non-state external threats, insider threats, and other expected threat actors, taking into account at least the current and expected capabilities of their models.” There is no numbered tier anywhere in that Measure — a signatory names its threat actors and then, under Measure 6.2, implements “appropriate security mitigations to meet the Security Goal,” staged as capabilities increase. The obligation is to be explicit about who you’re defending against, not to hit a numbered rung on someone else’s ladder.

Meta’s Advanced AI Scaling Framework, version 2, takes a related but distinct approach: it names two named risk-threshold tiers, High and Critical, and ties a specific measure to each rather than a numbered security level. Table 1 states that crossing the High threshold requires the developer to “initiate protocols for heightened access controls to model weights and security protections to prevent their tampering or exfiltration,” and crossing Critical requires the same access-control protocols “as overseen by the Chief AI Officer and/or Director of Alignment and Risk.” Meta’s document does not reference RAND or any external numbered scale anywhere in its security-mitigations language — “Heightened Access Controls” is Meta’s own term, applied at two thresholds, not five graded levels.

xAI’s approach is the hardest of the nine to verify in detail. Its safety document, a “Risk Management Framework (Draft)” dated February 20, 2025 and explicitly still marked as a draft, is corroborated as real by independent trackers (the Midas Project, the Future of Life Institute’s AI Safety Index), which describe it as identifying threat actors and mitigations without a named graded level — consistent with the goal-based camp rather than the scale-based one. The primary document itself could not be independently read for this piece: the PDF at its published URL currently returns an unresolved Git LFS pointer file rather than the actual document, a hosting-side issue on xAI’s own infrastructure rather than a takedown or access block. The Future of Life Institute’s Summer 2025 index scored xAI’s safety-framework category D+ and explicitly recommended xAI “boost current draft safety framework to match the efforts by Anthropic and OpenAI” — independent, if secondhand, confirmation that xAI’s framework was, as of that review, comparatively undeveloped.

G42’s numbered scale, built independently of RAND

G42’s Frontier AI Safety Framework takes the opposite structural choice from the goal-based camp while also not citing RAND: it defines its own numbered “Security Mitigation Levels” (SML), independent of any external benchmark named in the document. Security Level 1 is “suitable for models with minimal hazardous capabilities. No novel mitigations required.” Security Level 2 is described as “intermediate safeguards for models with capabilities requiring controlled access… secured such that it would be highly unlikely that a malicious individual or organization (state sponsored, organized crime, terrorist, etc.) could obtain the model weights or access sensitive data.” The framework ties SML directly to its own capability thresholds rather than an outside scale: “If a Frontier Capability Threshold has been reached, G42 will update this Framework to define a more advanced threshold that requires increased deployment (e.g., DML 3) and security mitigations (e.g., SML 3),” and states that “if a necessary Security Mitigation Level cannot be achieved, then further capabilities development of the model must be paused.” G42 is, on the available evidence, the one organization here with a genuinely independent numbered security scale — not adopted from RAND, and not phrased as a goal.

Microsoft: risk-tiered security scaling, without a distinctly named level of its own

Microsoft’s Frontier Governance Framework ties its security requirements to the same qualitative low/medium/high/critical capability-risk classification it uses throughout the document, rather than to a separately named “security level.” Its own text states: “Any model that triggers leading indicator assessment is subject to robust baseline security protection. Security safeguards are then scaled up depending on the model’s pre-mitigation scores, with more robust measures applied to models with High and Critical risk levels.” For models scoring High, the framework specifies concrete measures — restricted access-control-list hygiene, defense in depth across the lifecycle, and encrypted model weights among them.

NIKOLAI’s own crosswalk row for this element describes Microsoft’s approach as citing the RAND security-levels framework. The specific passage making that citation was not located directly in the sections of Microsoft’s own document available for this piece — worth flagging honestly rather than repeating the claim as independently confirmed. What is independently confirmed is that Microsoft’s public framework applies its own risk-tier labels (Low/Medium/High/Critical) to scale security measures, without introducing a separately branded numbered security-level term the way DeepMind or G42 do.

METR: security tied to capability, not to a fixed point in time

METR‘s December 2025 report Common Elements of Frontier AI Safety Policies, an analysis of the shared structure across twelve published labs’ frameworks, treats “Model Weight Security” as one of its named common elements rather than a single fixed bar: “Measures that will be taken to prevent model weight access by unauthorized actors. If malicious actors steal the weights of models with capabilities of concern… as models develop increasing capabilities of concern, progressively stronger information security measures are recommended to prevent theft and unintentional release of model weights.” That framing — security scaling with capability rather than being fixed once at a single named tier — is a description of the field as a whole, not a policy commitment of METR’s own; METR is an evaluator, not a model developer, and doesn’t operate a frontier model security program itself. It is, however, the closest thing to an independent, cross-organizational summary of how the other eight approaches actually relate to each other, and the RAND-citing detail above for both DeepMind and Magic’s AGI Readiness Policy (a developer outside NIKOLAI’s current nine-organization crosswalk) came from a direct reading of that same report.

The real divide: graded scale vs. named goal

Read across all nine, the split that matters isn’t RAND-adopters vs. everyone else — it’s whether an organization ranks its security posture on an ordinal scale at all.

Organization Their term Structure Cites RAND?
Google DeepMind Security Level 2+ / 3 / 4 Graded scale Yes, explicitly — adopts RAND’s SL levels directly
Anthropic ASL-3 Security Standard Graded scale (own ASL system) Only as an industry-wide recommendation, not its own commitment
G42 Security Mitigation Levels (SML) 1–4 Graded scale (independent) No — own numbered scale
Microsoft Risk-tiered security scaling (Low/Medium/High/Critical) Graded, but tied to capability-risk tiers, not a separate security scale NIKOLAI’s mapping says yes; not independently located in this piece’s research
OpenAI High standard / “a Critical standard” Two named tiers; Critical named but not yet specified No
xAI Security Goal Named goal, no ranking No
Meta Heightened Access Controls (Table 1) Named measure at two thresholds (High, Critical) No
EU GPAI Code Security Goal (Measure 6.1) Named goal, no ranking — explicit rejection of a graded model No
METR Model Weight Security Capability-contingent description, not a fixed bar Cites RAND when describing DeepMind’s and Magic’s approaches

Four organizations — DeepMind, Anthropic, G42, and (per NIKOLAI’s reading) Microsoft — use some form of ordinal scale. Three — xAI, Meta, and the EU’s Code — explicitly require a stated goal instead, with Meta’s two named thresholds sitting closer to the goal camp than the scale camp since “Heightened Access Controls” isn’t ranked against anything. OpenAI sits awkwardly in between: it names two tiers (High, Critical) the way a scale-based approach would, but only fully specifies the lower one. Only DeepMind and, per NIKOLAI’s shadow mapping, Microsoft explicitly anchor to the external RAND framework; G42’s scale is real but built independently of it.

What NIKOLAI maps, and what it doesn’t

CASRAI’s own NIKOLAI project — an independent, unendorsed reference dictionary for frontier AI safety, not a standard any lab or regulator has agreed to — formalizes Security Level as element N6: Security level, part of Track N6, Mitigations and Security. NIKOLAI’s own definition states: “NIKOLAI proposes Security level as a graded controlled-value statement of a developer’s model-weight and infrastructure security posture, indexed to a named external scale (e.g., RAND Security Levels) where the developer has adopted one. This represents an unsourced editorial synthesis; organizations vary in whether they use graded scales at all” — NIKOLAI’s own page is explicit that a graded, RAND-indexed structure is CASRAI’s proposed shape for the element, not a shape every organization actually uses, which is exactly what this piece found in the primary documents: three of nine reject a graded scale outright.

This is the first NIKOLAI element to get its own standalone deep-dive page rather than appearing only inside a broader comparison — a deliberate choice, since Security Level’s nine-way crosswalk is unusually rich evidence for the scale-vs-goal divide that runs through the whole N6 track (the same track also covers robustness level, coverage level, and mitigation change, each with its own crosswalk). Every row above — and every row in NIKOLAI’s own published crosswalk table for this element — is what NIKOLAI calls a shadow mapping: CASRAI’s own reading of a published document against NIKOLAI’s element definition, not a confirmation from the mapped organization that this is how it would describe itself. None of the nine organizations wrote their framework against NIKOLAI’s schema, and none has filed a formal Mapping Declaration accepting or correcting how NIKOLAI has read their document. Where this piece’s own research couldn’t independently confirm a specific claim in NIKOLAI’s crosswalk — Microsoft’s RAND citation, and the internal structure of xAI’s still-unreadable draft framework — that gap is flagged above rather than smoothed over.

For the companion element deep-dives already live, see NIKOLAI’s Capability Thresholds crosswalk (N3) and Evaluator Independence crosswalk (N8). For the full N1–N10 track structure, see NIKOLAI’s track system, and for what NIKOLAI is and isn’t, see What Is NIKOLAI?

Frequently asked questions

Do Google DeepMind and Anthropic use the same security scale?

Both reference the same external source — RAND’s report on securing AI model weights — but only DeepMind directly names RAND’s SL2, SL3, and SL4 levels as its own company commitments. Anthropic’s RSP ties its own ASL-3 Security Standard to Anthropic’s separate ASL system; RAND SL4 appears in the RSP only as an industry-wide recommendation for threat actors outside any governance regime, not as Anthropic’s own defined standard.

What does OpenAI’s “Critical standard” actually require?

As of the version reviewed for this piece, OpenAI’s Preparedness Framework does not specify one. Its own safeguard guidance for the Critical capability threshold states the organization will “halt further development” until a Critical security standard has been specified — the standard is named as a threshold to be met, but its content isn’t yet written down in the document.

Which organizations reject a graded security-level scale entirely?

xAI, Meta, and the EU AI Act’s GPAI Code of Practice each require a stated security goal — naming the threat actors a developer’s mitigations are meant to defend against — rather than a ranked numbered level. The EU Code’s Measure 6.1 states this requirement directly by name (“Security Goal”); xAI and Meta use the same structural approach without using that exact label.

Is G42’s Security Mitigation Levels scale based on RAND’s framework?

Not on the evidence available for this piece. G42’s own framework ties its Security Mitigation Levels (SML 1–4) to its own Frontier Capability Thresholds, without citing RAND or any other named external scale in the passages describing SML — distinct from DeepMind’s directly RAND-indexed approach.

Is NIKOLAI’s Security Level crosswalk an official standard?

No. NIKOLAI is CASRAI’s own, unendorsed reference project. Its crosswalk rows are shadow mappings — CASRAI’s own reading of each organization’s published documents against NIKOLAI’s element definition — not declarations the mapped organizations have made about themselves, and not a certification that any organization’s security posture meets any particular bar.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about Security Level: Nine Labs and Regulators, One Undefined Standard

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →