Source of record
Where this definition comes from
Meta Advanced AI Scaling Framework v2, Appendix I, p.41
“"Frontier AI" has two criteria: "High Capabilities in Catastrophic Risk Areas" and "Compute Threshold: We trained the model using at least 10^26 integer or floating point operations (to include material modifications ...)"”
https://ai.meta.com/static-resource/Meta_Advanced-AI-Scaling-Framework-v2California SB 53, 22757.11(i)
“"Frontier model": "a foundation model that was trained using a quantity of computing power greater than 10^26 integer or floating-point operations", including "any subsequent fine-tuning, reinforcement learning, or other material modifications"”
https://leginfo.legislature.ca.gov/faces/billTextClient.xhtml?bill_id=202520260SB53EU AI Act Article 51, Article 51(1)-(3)
“Article 51(1): "(a) it has high impact capabilities evaluated on the basis of appropriate technical tools and methodologies" or "(b)" a Commission designation decision; Article 51(2): presumed high-impact "when the cumulative amount of computation used for its training measured in floating point operations is greater than 10^25", rebuttable.”
https://ai-act-service-desk.ec.europa.eu/en/ai-act/article-51
Crosswalk
How named organisations use this concept
| Organisation | Their term, as published | Match | Source |
|---|---|---|---|
| Anthropic Advanced AI Framework | “Covered Developer test (>10^25 training FLOP or >$500M annual AI-derived revenue or >$1B AI R&D spend); and "Over time, the FLOP needed to train dangerous models may fall, and it may make sense to introduce a threshold based on capabilities rather than simply on training costs." (AAF p.3)” | close confidence: medium | Anthropic Advanced AI Framework, legislative proposal |
| OpenAI Frontier Governance Framework / Preparedness Framework v2 | “"frontier models as defined under the TFAIA and 'general-purpose models with systemic risk' under the EU AI Act" (FGF §1). PF: "Covered deployment": "any new or updated deployment that has a plausible chance of reaching a capability threshold whose corresponding risks are not addressed by an existing Safeguards Report" (§3.2).” | close confidence: medium | OpenAI Frontier Governance Framework |
| Google DeepMind Frontier Safety Framework v3.1 | “"Frontier AI Models or Models: are trained on a large data set, display significant generality, are capable of performing a wide range of distinctive tasks and have high-impact capabilities. Frontier AI models' agentic and reasoning-based general capabilities near or exceed those of other Google models." (glossary, p.18). This is a relative test, not an absolute number.” | close confidence: medium | Google DeepMind Frontier Safety Framework v3.1 |
| xAI Frontier AI Framework, 30 Jun 2026 | “"xAI's frontier AI models, such as Grok" (s.1); no compute or capability criterion given at all.” This source's PDF metadata /Title reads "Privileged/Confidential DRAFT working FRAMEWORK DOC"; no xAI statement disambiguating draft vs. final status was found (open [VERIFY] item in the source document's register). | none confidence: medium | xAI Frontier AI Framework, 30 Jun 2026 |
| Meta Advanced AI Scaling Framework v2 | “"Frontier AI" has two criteria: "High Capabilities in Catastrophic Risk Areas" and "Compute Threshold: We trained the model using at least 10^26 integer or floating point operations (to include material modifications ...)" (Appendix I, p.41).” | exact confidence: high | Meta Advanced AI Scaling Framework v2 |
| EU (AI Act / GPAI Code of Practice) AI Act Article 51 / Safety and Security chapter | “Article 51(1) sets two alternative tests: "(a) it has high impact capabilities evaluated on the basis of appropriate technical tools and methodologies, including indicators and benchmarks"; or "(b) based on a decision of the Commission ... having regard to the criteria set out in Annex XIII." Article 51(2): "A general-purpose AI model shall be presumed to have high impact capabilities pursuant to paragraph 1, point (a), when the cumulative amount of computation used for its training measured in floating point operations is greater than 10^25", rebuttable, with Article 51(3) empowering the Commission to amend the threshold by delegated act.” Open [VERIFY] item preserved from the source: no "10^23" figure appears anywhere in Article 51 -- a "10^23 GPAI indicator" cited in earlier drafts has not been located in this source and needs tracing to Annex XIII or the Commission's July 2025 Guidelines PDF before it can be repeated as an EU coverage-scope test. Article 3(65)'s own definitional wording is still known only via the chapter's Appendix 1.2 paraphrase, not read directly, which is also [VERIFY]. Secondary citation tag {DEU} in the source document does not resolve to a URL in the provided sources table. | exact confidence: high | EU AI Act Article 51 |
| California SB 53 SB 53 statute text | “"Frontier model": "a foundation model that was trained using a quantity of computing power greater than 10^26 integer or floating-point operations", including "any subsequent fine-tuning, reinforcement learning, or other material modifications" (22757.11(i)).” | exact confidence: high | California SB 53 |
| US Government (EO 14409) Executive Order 14409 | “"a classified benchmarking process to assess the advanced cyber capabilities of AI models and determine the threshold at which an AI model should be designated a 'covered frontier model'" (Sec. 3(a)). Cyber-only and classified.” | narrow confidence: medium | Executive Order 14409 |
| Demis Hassabis (personal framework essay) "A framework for frontier AI" essay | “"A model would qualify as 'Frontier-class' if it meets certain thresholds on a set of benchmarks determined by the Standards Body and regularly updated" (§Framework ¶2).” Personal essay, not an official DeepMind policy document. | close confidence: medium | Demis Hassabis, "A framework for frontier AI and the dawning of a new age" |
| Microsoft Frontier Governance Framework (Microsoft) | “Scopes its leading-indicator assessment to models in scope under the EU AI Act, TFAIA and RAISE, and to substantial fine-tunes over one-third of base compute (para.).” | close confidence: medium | Microsoft Frontier Governance Framework |
| US Congress (H.R. 9925, FRONTIER Act, not enacted) FRONTIER Act bill text | “Sets a tiered test: "frontier developer" at >10^26 operations; "large frontier developer" additionally requires >$50M revenue and >=$1B in AI-related development expenditure over a rolling 36-month window; "very large frontier developer" additionally requires >$5B revenue and >=$10B in AI spending -- each tier inheriting the duties of the tiers below it (Sec. 2, Sec. 6).” Bill not enacted. | close confidence: medium | FRONTIER Act, 119th Congress (H.R. 9925, not enacted) |
Divergence
Where sources materially disagree
"Frontier model" is itself a false-friend label across this corpus (see source §4): 10^26 FLOP (SB 53, Meta); 10^25 FLOP plus revenue (Anthropic AAF); relative to Google's own models (GDM); a classified cyber benchmark (EO 14409); a benchmark set refreshed by a standards body (Hassabis). Critically, the EU AI Act itself never uses the phrase "frontier model" at all -- its gate is the distinct "general-purpose AI model with systemic risk" classification under Article 51, with its own two-part test. NIKOLAI should treat "frontier model" and "GPAI model with systemic risk" as related but non-identical scope concepts, not as translations of one another. "Systemic risk" is also a term of art specific to the EU chapter and is not the same scope as SB 53's "catastrophic risk" (50 deaths/$1B, single-incident) or the labs' "severe harm"; xAI's FAIF borrows the EU's four-domain taxonomy but not its legal test.
Gap
Six incompatible scope tests are in use (10^25 FLOP plus revenue; 10^26 FLOP; relative to the developer's own models; a classified cyber benchmark; a benchmark set that a standards body would refresh; and H.R. 9925's three-tier compute-plus-revenue-plus-spending ladder). A risk report cannot be compared across regimes without recording which test it was produced under.







