Written and maintained by CASRAI Editorial Board
Last updated
Last verified: September 20, 2026. The European Union’s AI Act never uses the phrase “frontier model.” Not once, anywhere in Article 51. Its gate is a different legal construct entirely — a “general-purpose AI model with systemic risk” — with its own two-part test and its own vocabulary. Meanwhile California’s SB 53 defines “frontier model” in one sentence and “catastrophic risk” in another, and the two are not synonyms for the EU’s “systemic risk” or for the “severe harm” language labs use internally. Four terms, three of them genuinely different legal or technical objects, all describing roughly the same governance moment — the point where a model becomes big enough, or capable enough, that a framework’s obligations switch on. This guide is a deep-dive on that switch: the if-then test, across eleven organizations, that decides whether a given developer or model is in scope at all.
The short version, stated plainly: no two of the eleven tests below are identical. Three use an absolute compute bar (Meta, SB 53, the EU’s rebuttable presumption) but set it at two different numbers. Two combine compute with a financial-scale gate, differently structured (Anthropic’s AND-plus-OR; the FRONTIER Act’s three-tier ladder). One test is relative, not absolute (Google DeepMind, measured against Google’s own prior models). One is classified (Executive Order 14409). One is pure self-identification with no criterion at all (xAI). Two scope themselves to other regimes’ definitions rather than writing their own (OpenAI, Microsoft). One is a named individual’s personal proposal for a benchmark a standards body hasn’t been created yet to set (Demis Hassabis). A risk report, a compliance filing, or a piece of journalism that says a model is or isn’t “in scope” without saying which of these eleven tests it means is not actually saying anything comparable across frameworks.
What a Coverage Scope Threshold Is Supposed to Do
Structurally, every test below does the same job: it is the condition that decides whether a given developer or a given model falls under a framework’s obligations at all — before any question of capability thresholds, safety cases, or incident reporting even arises. Get the scope test wrong and every downstream obligation in a framework is moot for that model. The tests differ on four independent design choices: whether the trigger is the developer (a company-level test, like Anthropic’s or the FRONTIER Act’s) or the model (a model-level test, like Meta’s or SB 53’s); whether the number is absolute (a fixed FLOP count) or relative (compared to another model, as with Google DeepMind); whether crossing the line is automatic or a rebuttable presumption subject to override (the EU’s is explicitly rebuttable; most others are not); and whether the criterion is public at all (Executive Order 14409’s is classified by design). Reading the eleven rows side by side makes clear that “in scope” is not one idea wearing different words — it is at least six structurally different ideas that happen to share a family resemblance.
The Labs: Two Absolute Bars, One Relative Test, One Combined Test, and One Non-Test
Meta — Either Gate Trips It, Not Both
Meta’s Advanced AI Scaling Framework v2 (Appendix I, p.41) sets two alternative criteria for “Frontier AI” status, joined by or, not and: “High Capabilities in Catastrophic Risk Areas” — a model “reasonably likely to be more capable” in a catastrophic risk domain (Chemical & Biological, Cybersecurity, or Loss of Control) than Meta’s existing models or the most advanced models generally — or a “Compute Threshold: We trained the model using at least 1026 integer or floating point operations,” a figure that explicitly “include[s] material modifications to the model through fine-tuning, reinforcement learning training, and other training steps.” Either gate alone is sufficient. A model that clears 1026 FLOP is in scope even if nothing about its measured capabilities looks unusual, and a model well under that compute figure is in scope anyway if its capabilities in a catastrophic risk domain are judged comparable to Meta’s existing frontier models.
Anthropic — A Compute Floor, Plus One of Two Financial Gates (Not a Flat Three-Way Or)
Anthropic’s Advanced AI Framework, its legislative proposal for a federal frontier-AI statute, states the “Covered Developer” test as two criteria joined by both, not a flat three-way or: “The proposed obligations apply to AI developers that meet both of the following criteria: Develop AI models requiring more than 1025 training FLOP” and “Earn more than $500 million in annual AI-derived revenue, or spend more than $1 billion per year on AI research and development.” In other words: the compute floor is mandatory, and it is paired with whichever of the two financial-scale gates the developer clears. A developer under the compute floor is out of scope regardless of revenue; a developer over the compute floor but under both financial gates is also out of scope. The framework separately flags that “[o]ver time, the FLOP needed to train dangerous models may fall, and it may make sense to introduce a threshold based on capabilities rather than simply on training costs” — an acknowledgment, from inside the proposal itself, that a pure compute gate has a shelf life.
OpenAI — Scoped to Other Regimes’ Definitions, Not Its Own
OpenAI’s Frontier Governance Framework does not set an independent compute or capability bar for what counts as in scope. Instead it borrows the tests other regimes already wrote: under California’s Transparency in Frontier AI Act, “this FGF is our Frontier AI Framework, documenting OpenAI’s technical and organizational protocols to manage, assess, and mitigate catastrophic risks, as defined under the TFAIA,” and under the EU’s GPAI Code of Practice, “this FGF serves as our publicly available summary of OpenAI’s Safety & Security Framework, describing how we assess and mitigate systemic risks… for models covered under Regulation (EU) 2024/1689 (the EU AI Act).” OpenAI’s own scope, in other words, is a function of California’s and the EU’s scopes — whichever models those two regimes already treat as in scope are the models the FGF covers.
Google DeepMind — Relative to Google’s Own Models, Not an Absolute Number
DeepMind’s Frontier Safety Framework v3.1 glossary defines the term without a FLOP figure at all: “Frontier AI Models or Models: are trained on a large data set, display significant generality, are capable of performing a wide range of distinctive tasks and have high-impact capabilities. Frontier AI models’ agentic and reasoning-based general capabilities near or exceed those of other Google models.” That last clause is the whole test — it is relative to Google’s own model lineage, not to an absolute compute or capability bar any outside party could apply independently. Two organizations training at identical compute and capability levels could land on opposite sides of DeepMind’s line depending entirely on what Google’s own most recent models already do.
xAI — Pure Self-Identification, No Criterion At All
xAI’s Frontier AI Framework (30 June 2026) states its scope in a single clause with no test attached: “This Frontier AI Framework (‘FAIF’) outlines xAI’s approach to policies for mitigation of significant risks associated with the development, deployment, and release of xAI’s frontier AI models, such as Grok” (§1). No compute figure, no capability comparator, no revenue gate — a model is in scope because xAI calls it Grok, or otherwise designates it frontier, not because it clears any published, independently checkable line. Worth flagging: the source PDF’s own metadata title reads “Privileged/Confidential DRAFT working FRAMEWORK DOC,” and no public xAI statement was found disambiguating draft from final status for this version.
The Regulators: A Rebuttable Presumption, a Hard Number, and a Classified Benchmark
EU AI Act Article 51 — “Systemic Risk,” Never “Frontier Model”
Article 51(1) sets two alternative tests for whether a general-purpose AI model is classified as having systemic risk: “(a) it has high impact capabilities evaluated on the basis of appropriate technical tools and methodologies, including indicators and benchmarks”, or “(b)” a designation decision by the European Commission, applying the criteria in Annex XIII. Article 51(2) then supplies a compute-based shortcut: a model is presumed to have high impact capabilities — not automatically classified, presumed, which is rebuttable — “when the cumulative amount of computation used for its training measured in floating point operations is greater than 1025.” Two things about this test are easy to get wrong. First, the word “frontier” does not appear in Article 51 — the EU’s legal category is “general-purpose AI model with systemic risk,” a term of art specific to this chapter, not a synonym or translation of “frontier model” as SB 53 or the labs use it. Second, an earlier drafting stage of the Act is sometimes cited as having used a lower “1023” compute indicator for GPAI models generally; that figure does not appear anywhere in the current Article 51 or Annex XIII text, and this guide does not repeat it as current EU law — it should be treated as an open item requiring its own sourcing before anyone cites it as live, not as a fact. See CASRAI’s fuller explainer on GPAI systemic risk under the EU AI Act for the Article 55 obligations that follow once a model clears this test.
California SB 53 — A Hard Number, Explicitly Retroactive to Fine-Tuning
SB 53 defines “frontier model” in a single sentence, at §22757.11(i): “a foundation model that was trained using a quantity of computing power greater than 1026 integer or floating-point operations,” and the statute is explicit that this includes “any subsequent fine-tuning, reinforcement learning, or other material modifications” to the original training run — a model cannot be engineered out of scope by fine-tuning it below the line after an initial training run that cleared it. SB 53 separately defines “catastrophic risk” at §22757.11(c) and §22757.16: “a foreseeable and material risk that a frontier developer’s development, storage, use, or deployment of a frontier model will materially contribute to the death of, or serious injury to, more than 50 people or more than one billion dollars ($1,000,000,000) in damage to, or loss of, property arising from a single incident” — and the statute separately clarifies that “the loss of value of equity does not count as damage to or loss of property.” That figure — 50 deaths or $1B in a single incident — is SB 53’s own defined term. It is not the same object as the EU’s “systemic risk” classification (a scope test, not a harm-magnitude test) or as the qualitative “severe harm” language labs use internally in their own safety frameworks. See CASRAI’s full SB 53 explainer for the transparency-report and incident-reporting obligations that attach once a developer and model clear these two tests.
US Executive Order 14409 — Classified, and Cyber-Only
Executive Order 14409 takes a fundamentally different approach to scope: it directs the government to “develop and maintain a classified benchmarking process to assess the advanced cyber capabilities of AI models and determine the threshold at which an AI model should be designated a ‘covered frontier model’” (Sec. 3(a)) — with that determination made by the NSA Director in consultation with other officials. The Order narrows its own scope twice over: the benchmark is cyber-capability-specific, not a general frontier-AI test, and its actual value is classified by design, unavailable for public comparison against any of the other ten tests in this piece. Section 3(b) separately establishes a voluntary channel for developers to engage federal agencies about whether their models meet “covered frontier model” status and to grant the government pre-release access; Section 3(c) explicitly disclaims any mandatory licensing or preclearance requirement.
Microsoft — Scoped to Everyone Else’s Definitions, Plus Its Own Fine-Tune Rule
Microsoft’s Frontier Governance Framework (February 2026) takes the same borrowed-scope approach as OpenAI, but names its sources explicitly and adds one original rule of its own: “A leading indicator assessment is run on any model that is in scope for frontier model requirements under applicable laws, such as the EU AI Act, California’s Transparency in Frontier AI Act (TFAIA), and New York’s Responsible AI Safety and Education (RAISE) Act.” The original addition is a fine-tuning trigger: “We also run a leading indicator assessment when Microsoft substantially fine-tunes first- or third-party models, where the compute used for fine-tuning is more than 1/3 of the base model’s” training compute — and where a third-party model’s base compute is unknown, Microsoft defaults to “1/3 of 1025 FLOPs as the threshold for substantial fine-tuning of third-party frontier models.” So Microsoft’s scope is, in effect, the union of three external legal definitions (EU AI Act, TFAIA, RAISE Act) plus one internally-authored fine-tuning rule layered on top.
H.R. 9925, the FRONTIER Act (Not Enacted) — A Three-Tier Compute-Plus-Spend Ladder
H.R. 9925, introduced in the US House and not enacted, proposes the most granular tiered scope test in this set — three developer classes, each nested inside the one before it, all sharing the same 1026 FLOP compute floor:
- Frontier developer: has trained, or initiated training of, a frontier model — a “foundation model that was trained using a quantity of computing power greater than 1026 integer or floating-point operations” — with no revenue or R&D-spending requirement.
- Large frontier developer: a frontier developer that, over the preceding 36-month period, “had gross revenues in excess of $50,000,000” and “incurred not less than $1,000,000,000 in AI-related development expenditures.”
- Very large frontier developer: a frontier developer that, over the same 36-month window, “had gross revenues in excess of $5,000,000,000” and “incurred not less than $10,000,000,000 in AI-related development expenditures.”
Because the bill has not been enacted, none of this is currently binding — it is included here as a proposed test, not a live one, and should be read that way.
One Outlier: A Standards Body That Doesn’t Exist Yet
Google DeepMind CEO Demis Hassabis’s personal essay, “A Framework for Frontier AI,” proposes a scope test built around an institution the essay itself acknowledges does not yet exist: “A model would qualify as ‘Frontier-class’ if it meets certain thresholds on a set of benchmarks determined by the Standards Body and regularly updated” to keep pace with advancing capabilities. This is worth separating clearly from DeepMind’s own Frontier Safety Framework, covered above — the essay is Hassabis’s personal proposal for how an industry-wide standards body might eventually set a shared scope test, not an official DeepMind policy document, and no such Standards Body currently exists to do the determining.
The Vocabulary Trap: Four Words That Are Not Interchangeable
Running through all eleven rows above is a terminology problem that matters more than any single number: “frontier model,” “systemic risk,” “catastrophic risk,” and “severe harm” get used as if they were four names for the same idea, and they are not.
“Frontier model” is the scope test itself — the question of which developers or models a framework’s obligations even apply to. It shows up in Meta’s and SB 53’s compute-threshold language and in most lab frameworks’ own vocabulary. It is absent from the EU AI Act entirely. “Systemic risk” is the EU’s own distinct legal category under Article 51 — a GPAI-model classification, not a harm-magnitude test, and not the EU’s version of “frontier model.” “Catastrophic risk” is SB 53’s defined harm-magnitude term — the specific 50-deaths-or-$1B-in-a-single-incident bar at §22757.11(c) — and Anthropic’s Advanced AI Framework borrows nearly identical language for its own “Catastrophic Risk” definition, but the EU chapter does not use this term at all. “Severe harm” is the qualitative language labs use internally across their own safety frameworks (OpenAI’s Preparedness Framework, for instance, ties its capability thresholds to risks that could “meaningfully increase risk of severe harm”) — a deliberately non-quantified standard, structurally different from SB 53’s numeric bar even when the two are gesturing at similar underlying concerns. Treating any two of these four terms as synonyms — especially “frontier model” and “systemic risk,” the single most common false-friend pairing in coverage of this space — produces a sentence that sounds precise and is not.
What the Eleven Rows Add Up To
Lay the eleven tests side by side and at least six structurally incompatible scope mechanisms are in active use: an absolute compute bar at 1026 FLOP (Meta, SB 53, the FRONTIER Act’s base tier); a lower absolute compute bar at 1025 FLOP, but only as a rebuttable presumption rather than an automatic classification (the EU); a compute floor combined with a financial-scale gate, structured two different ways (Anthropic’s AND-then-OR; the FRONTIER Act’s nested revenue-and-spending tiers); a relative test with no absolute number at all (Google DeepMind, measured against Google’s own models); a classified, cyber-domain-only benchmark (Executive Order 14409); and pure self-identification with no published criterion whatsoever (xAI). Two frameworks (OpenAI, Microsoft) opt out of writing an original scope test altogether and instead inherit whatever the EU, California, and New York have already defined. A risk report, a piece of journalism, or a compliance filing that asserts a model is “in scope” without naming which of these six mechanisms it means has not actually said anything comparable across frameworks — and the closer two organizations’ vocabulary sounds (“frontier model” versus “systemic risk” being the clearest case), the more that gap tends to go unnoticed.
Where NIKOLAI Fits In
This page is a direct deep-dive on NIKOLAI’s coverage-scope-threshold element, part of Track N1, Actors, Models and Scope, in CASRAI’s frontier-AI-safety dictionary — and it opens N1’s first per-element crosswalk page, joining Track N8’s evaluator-independence deep-dive and Track N3’s capability-threshold deep-dive as the structural pattern this piece follows: one NIKOLAI element, every organization’s published language on it, laid out side by side. Every row above reflects only what that organization has explicitly published; NIKOLAI itself labels every one of these eleven rows a “shadow mapping” unless the organization has gone through NIKOLAI’s own Mapping Declarations process to confirm how the term maps to its usage. As of this writing, none of the eleven has made that declaration, so every mapping above is CASRAI’s own independent read of the public record — not a confirmation from Meta, Anthropic, OpenAI, Google DeepMind, xAI, the EU, California, the US government, Microsoft, the FRONTIER Act’s sponsors, or Demis Hassabis. NIKOLAI is CASRAI’s own project: an independent, unendorsed reference work, not an official standard adopted by any framework it maps.
NIKOLAI’s own gap note on this element counts six incompatible scope tests in active use — the same six mechanisms this guide walks through above — and proposes that a risk report or transparency filing record which specific test a coverage claim was produced under, as a structured value with a magnitude, a unit, and any stated exclusions, rather than treating “in scope” as a single free-text boolean. That is the same diagnosis this guide reaches independently: the fix is not agreeing on one universal scope test, which eleven governments and companies show no sign of converging on, but making explicit, every time, which of the six is being invoked.
A Note on Scope: What This Page Does and Doesn’t Cover
This page is deliberately narrower than three other CASRAI guides it overlaps with, and non-duplicative of each. What Is a Frontier AI Model? introduces the three definitional approaches (compute, capability, relative state-of-the-art) at an explainer level for a reader encountering the concept for the first time; this page assumes that context and instead runs the full eleven-organization crosswalk with exact quoted figures and section citations. GPAI Systemic Risk: The EU AI Act Term Explained covers Article 51’s classification test and the Article 55 obligations that follow it in much greater depth than the one section devoted to it here. California SB 53: The Foundational Explainer covers SB 53’s full transparency-report, incident-reporting, and whistleblower-protection regime; this page uses only SB 53’s two scope-defining terms. Read this page for the cross-framework comparison; read those three for the full mechanics of any one regime.
Not Yet Mapped: Singapore and India
Two jurisdictions with active AI-governance discussion — Singapore and India — do not currently have a published, binding coverage-scope test comparable to the eleven above, and CASRAI has not published guides on either as of this writing. They are flagged here as candidate rows for a future update to NIKOLAI’s coverage-scope-threshold crosswalk, not as existing entries; nothing above should be read as implying either jurisdiction has a mapped scope test today. See CASRAI’s jurisdiction-by-jurisdiction map of AI regulation for the broader multi-jurisdiction picture this element sits inside.
Frequently Asked Questions
Does the EU AI Act use the term “frontier model”?
No. Article 51 never uses the phrase. Its legal category is “general-purpose AI model with systemic risk,” classified under a two-part test in Article 51(1)-(2) — a distinct construct from “frontier model” as SB 53 or the labs define it, even though both are trying to identify roughly the same category of consequential model.
Is “systemic risk” the same as “catastrophic risk”?
No. “Systemic risk” is the EU AI Act’s Article 51 classification category for a GPAI model. “Catastrophic risk” is California SB 53’s defined harm-magnitude term, set at more than 50 deaths or more than $1 billion in property damage or loss arising from a single incident (§22757.11(c); §22757.16). They come from different legal instruments and measure different things.
Does a 1023 FLOP threshold appear anywhere in the current EU AI Act text?
Not in Article 51 or Annex XIII as currently in force. A lower figure is sometimes cited from earlier drafting-stage discussion, but this guide could not locate it in the current text and does not repeat it as live law. Article 51(2)’s operative figure is 1025 FLOP, as a rebuttable presumption.
Is Anthropic’s Covered Developer test an “or” across all three criteria?
No. Anthropic’s Advanced AI Framework requires the compute floor (more than 1025 training FLOP) as a mandatory condition, combined with either of two financial-scale gates (more than $500 million in annual AI-derived revenue, or more than $1 billion in annual AI R&D spend). A developer under the compute floor is out of scope regardless of revenue.
Which of the eleven organizations has published the highest compute bar?
Meta, California SB 53, and the FRONTIER Act’s base “frontier developer” tier all use the same figure: greater than 1026 FLOP. The EU AI Act’s rebuttable presumption is lower, at 1025 FLOP; Anthropic’s compute floor is also 1025 FLOP but is paired with a mandatory financial-scale gate.
Has any of these eleven organizations confirmed NIKOLAI’s mapping of their scope test?
No. As of this writing, none of the eleven has gone through NIKOLAI’s Mapping Declarations process. Every row in this crosswalk is CASRAI’s own independent read of each organization’s published text — a “shadow mapping” in NIKOLAI’s own terminology, not a confirmation from the organization itself.







