Written and maintained by CASRAI Editorial Board
Last updated
Last verified: September 25, 2026. Frontier AI governance rests on a load-bearing assumption almost nobody states out loud: that you can see a dangerous capability coming before it arrives. Capability thresholds, early-warning evaluations, staged deployment, pre-registration of safety cases — every one of these mechanisms assumes that capability moves at a pace you can measure and project. The research literature on “emergent abilities” is the place where that assumption gets tested, and the answer it gives is more useful than either side of the headline debate suggests.
The short version: the strongest claim — that capability jumps are inherently unforecastable surprises — did not survive contact with the evidence. But the reassuring counter-claim, that emergence is purely a measurement artifact and therefore everything is predictable, does not survive either. What actually survives is narrower and more operationally awkward: aggregate performance trends are broadly forecastable, specific benchmark scores are forecastable only with effort and only a short distance ahead, and the thing governance frameworks actually need to forecast — whether a named threshold will be crossed before the next scheduled evaluation — is the hardest case of all.
Where the Claim Started: Wei et al., 2022
The term entered the field with “Emergent Abilities of Large Language Models” (Jason Wei, Yi Tay, Rishi Bommasani, Colin Raffel, Barret Zoph, Sebastian Borgeaud, Dani Yogatama, Maarten Bosma, Denny Zhou, Donald Metzler, Ed H. Chi, Tatsunori Hashimoto, Oriol Vinyals, Percy Liang, Jeff Dean and William Fedus), posted to arXiv on 15 June 2022 and published in Transactions on Machine Learning Research. Its definition is deliberately spare: an ability is emergent “if it is not present in smaller models but is present in larger models.”
Two properties made this interesting to anyone thinking about risk. The first is sharpness — performance appears to move from near-chance to substantial in a narrow band of scale. The second is unpredictability — the band appears at a scale you could not have named in advance by extrapolating from smaller models. The paper’s closing implication was that further scaling would unlock further abilities nobody had forecast.
That framing was already in tension with an observation Anthropic researchers had published a few months earlier. “Predictability and Surprise in Large Generative Models” (Deep Ganguli, Danny Hernandez, Liane Lovitt and colleagues), presented at ACM FAccT 2022, described the paradox directly: training loss follows smooth scaling laws across the training distribution, while the specific capabilities and specific failure modes that fall out of that loss remain hard to anticipate. Predictability at the aggregate level was fuelling investment; unpredictability at the capability level was leaving the harms unanticipated. That gap, not the word “emergent,” is the governance problem.
The Rebuttal: Metric Choice, Not Model Behaviour
In April 2023 Rylan Schaeffer, Brando Miranda and Sanmi Koyejo posted “Are Emergent Abilities of Large Language Models a Mirage?” It won an outstanding paper award at NeurIPS 2023, and it reframed the debate.
Their argument is a measurement argument, not a modelling one. For a fixed set of model outputs, they contend, apparent emergence arises from the researcher’s choice of metric rather than from any discontinuity in the underlying model. Nonlinear or discontinuous metrics manufacture sharp transitions; linear or continuous metrics over the same outputs show smooth, predictable improvement. Exact-string-match on a multi-token answer is the canonical offender: a model steadily getting more of the digits right scores zero until it gets all of them right, at which point it scores one. Multiple-choice grading behaves similarly — a steadily rising probability on the correct option shows nothing until it overtakes the alternatives.
The paper backs this three ways: predictions about metric-dependent emergence tested on the InstructGPT and GPT-3 family, a meta-analysis of which BIG-Bench tasks show emergence under which metrics, and — the most persuasive step — deliberately inducing apparently emergent behaviour in ordinary vision networks purely by changing the metric applied to fixed outputs.
What the Rebuttal Does and Does Not Settle
It settles this: sharpness is often an artifact. A graph with a hockey stick in it is not, on its own, evidence of a discontinuity in the model.
It does not settle unpredictability, and the distinction matters enormously for governance. Three limits are worth being precise about.
- The smooth metrics were found afterwards. Knowing that some continuous reparameterisation of a task shows smooth scaling is not the same as being able to name that reparameterisation before the jump happens. The mirage analysis is largely retrospective.
- Some governance-relevant metrics are irreducibly discontinuous. Whether a model can complete an end-to-end attack chain, or produce a working exploit, or finish a multi-step task without human correction, is a pass/fail question by nature. Smoothing it produces a number that is easier to forecast and harder to act on. A threshold is a step function; you cannot regulate the step away.
- Downstream prediction stayed hard even after the metric critique landed. Schaeffer and a larger group — including Hailey Schoelkopf, Stella Biderman and Koyejo — returned to this in “Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?” (June 2024). Their diagnosis: downstream scores are computed from log-likelihoods through a chain of transformations that progressively degrades the statistical relationship with scale. On multiple-choice benchmarks, predicting the score requires predicting not just how probability mass concentrates on the correct option but how it fluctuates across the incorrect ones. The same authors who showed emergence could be smoothed away also showed why smoothing it does not automatically buy you forecasts.
A third strand offers a mechanism rather than a metric. Tung-Yu Wu and Pei-Yu Lo’s “U-shaped and Inverted-U Scaling behind Emergent Abilities of Large Language Models” (ICLR 2025) splits benchmark questions by difficulty and finds opposing trends: hard questions scale U-shaped, easy questions inverted-U then upward. While the two cancel, aggregate performance looks flat; when the easy-question trend reverts to standard scaling, the aggregate appears to leap. They propose a forecasting pipeline, Slice-and-Sandwich, that exploits the difficulty split to predict both the emergence threshold and performance past it. Note what this implies: the jump is real in the aggregate and predictable — but only if you disaggregate first.
What Can Actually Be Forecast Today
Five approaches have demonstrated something concrete. Their demonstrated reach is the column that matters for governance.
| Approach | What it predicts | Demonstrated reach | Main limitation |
|---|---|---|---|
| Pretraining-loss scaling laws (GPT-4 Technical Report, March 2023) | Final training loss; pass rate on a HumanEval subset | Loss predicted from models using at most 10,000× less compute; HumanEval pass rate from at most 1,000× less compute | The report itself states certain capabilities remain hard to predict; the Inverse Scaling Prize task Hindsight Neglect is given as a case where GPT-4 reversed the prior trend |
| Observational scaling laws (Yangjun Ruan, Chris J. Maddison, Tatsunori Hashimoto, May 2024) | Downstream benchmark performance across model families, without training new models | Fit across roughly 100 publicly available models; several emergent phenomena shown to follow smooth sigmoidal curves predictable from small models, including agent performance and chain-of-thought effects | Requires a broad public population of released models; treats families as differing only in compute-to-capability efficiency |
| Emergence laws via finetuning (Charlie Snell, Eric Wallace, Dan Klein, Sergey Levine, November 2024) | Whether a future model will emerge on a task where current models score at chance | Validated on MMLU, GSM8K, CommonsenseQA and CoLA; in some cases predicted emergence for models trained with up to 4× more compute | “In some cases” is the operative phrase; 4× compute is a short horizon relative to release cycles |
| Difficulty-sliced forecasting / Slice-and-Sandwich (Tung-Yu Wu, Pei-Yu Lo, ICLR 2025) | The emergence threshold and post-threshold performance | Demonstrated on benchmarks exhibiting U-shaped / inverted-U difficulty splits | Needs per-question difficulty labelling; depends on the difficulty split being available and stable |
| Agent-benchmark forecasting (Govind Pimpale, Axel Højmark, Jérémy Scheurer, Marius Hobbhahn, February 2025) | Frontier agent scores on safety-relevant benchmarks | Six methods compared on 38 models; two-step method applied to SWE-Bench Verified, Cybench and RE-Bench, forecasting 54% (non-specialised) and 87% (state-of-the-art) on SWE-Bench Verified by early 2026 | The authors flag that their approach may be too conservative given inference-compute scaling |
Read together: capability forecasting is a real and improving discipline, not a category error. It is also nowhere near the horizon that a threshold-triggered governance regime implicitly assumes. A method that reliably calls emergence 4× compute ahead is useful. It is not the same as knowing, at the start of a training run, whether the finished model will cross a bright line.
How Safety Frameworks Compensate
Frontier safety frameworks do not resolve the forecasting problem. They engineer around it, in three recognisable moves.
1. Evaluate often enough that you cannot miss the crossing
Google DeepMind’s Frontier Safety Framework commits to developing evaluation suites it calls early warning evaluations, which will “alert us when a model is approaching a CCL,” run “frequently enough that we have notice before that threshold is reached.” That phrasing is a direct admission: because you cannot forecast the crossing, you sample densely enough that you catch it in progress. The cadence, not the forecast, is doing the safety work.
Anthropic’s Responsible Scaling Policy makes the same trade explicit and has adjusted it in public. Version 3.0, effective 24 February 2026, “clarifies the ambiguous definition of the evaluation interval, and extends the interval to 6 months” — lengthening the gap specifically to avoid lower-quality, rushed capability elicitation. The current policy at the time of writing is version 3.4, effective 8 July 2026. Anthropic’s own account of earlier practice is candid about the epistemic footing: it “relied on empirical observations and rough predictions” for scaling buffers, and notes that quantitative predictions “may become necessary” in future. A longer interval buys better evaluations at the cost of a wider window in which an unforecast crossing can occur; there is no setting of that dial that avoids the trade.
2. Act on the forecast, not only the measurement
OpenAI’s Preparedness Framework version 2 (15 April 2025) applies safeguards to models that have reached or are forecasted to reach a Critical capability level in a Tracked Category, requiring security and safety controls during development regardless of whether or when the model is externally deployed. This is the honest structural answer: if you cannot know, act as though the forecast were the measurement. It also quietly makes the forecast a compliance artifact. A developer’s internal capability projection stops being a research output and becomes the thing that triggers, or fails to trigger, a control.
3. Set the threshold with a margin, and accept the cost
Every framework that names a threshold is implicitly naming a margin: the gap between the trigger level and the level at which harm actually becomes plausible. Wider margins tolerate worse forecasts and cost more in unnecessary safeguards. METR’s Common Elements of Frontier AI Safety Policies (16 December 2025) documents how differently developers set that gap, and how much of the variation is undisclosed. If you want the term-level detail, our guide on how fourteen labs and regulators use “capability threshold” without a shared definition shows how little the published language constrains anyone. METR — Model Evaluation and Threat Research — is the evaluator most often cited on this point.
Why Regulators Reach for Compute Instead
The clearest regulatory response to unforecastable capability is to stop trying to forecast capability. The EU AI Act does exactly this. Article 51(1) classifies a general-purpose AI model as carrying systemic risk if it has high-impact capabilities, evaluated “on the basis of appropriate technical tools and methodologies, including indicators and benchmarks,” or by Commission decision. Article 51(2) then supplies the operative shortcut: a model “shall be presumed to have high impact capabilities” when cumulative training computation exceeds 1025 floating point operations. Article 51(3) empowers the Commission to amend the thresholds by delegated act so they “reflect the state of the art.”
Compute is a proxy that is observable before the capability exists, auditable after the fact, and immune to metric disputes. It is also a bad predictor of any particular ability — which is precisely why Article 51(2) is drafted as a rebuttable presumption sitting underneath a capability test, and why Article 51(3) exists at all. The amendment power is an acknowledgement that the proxy will drift as algorithms improve. See our guide on FLOP thresholds in compute governance and export controls for how the same proxy is used on the security side.
Where This Touches Research Administration
Three parts of a university research office inherit this problem whether or not anyone there reads scaling-law papers.
Research computing. Compute thresholds do not carve out academic training runs. An institution that aggregates GPU capacity for a large training or continued-pretraining project needs a defensible record of cumulative training FLOP — not because anyone expects a campus cluster to approach 1025, but because the number is the thing a regulator or funder will ask for, and it is far cheaper to log during the run than to reconstruct afterwards.
Research security and export control. Model weights and training methods can fall within export-control scope, and release decisions for open-weight artifacts turn on capability assessments that the forecasting literature says are uncertain. A research security office reviewing an open-weight release cannot resolve that uncertainty; it can require that the capability claim in the release memo be sourced, dated, and tied to a named evaluation rather than asserted.
Sponsored programs and institutional AI policy. Institutional policies that gate a tool on its capability level — “approved for use where the model cannot X” — are threshold policies with the same forecasting weakness as a frontier framework, minus the evaluation budget. The practical mitigation is the same one the labs use: fix a review cadence rather than trusting the capability claim to stay true.
Where NIKOLAI Fits
CASRAI maintains NIKOLAI, its own independent frontier-AI-safety dictionary. It is unendorsed: crosswalk rows are shadow mappings drawn from published documents, not positions any organisation has agreed to, unless that organisation has filed a Mapping Declaration. Two elements bear directly on this topic.
- Capability threshold (track N3, thresholds and checkpoints) records a level of model capability — or capability plus usage — at which specified additional safeguards or decisions become required, together with its disclosure status: quantified, qualitative, referenced-but-undefined, or classified. The forecasting literature explains why so many published thresholds are qualitative. If you cannot reliably predict a benchmark score two model generations out, committing to a number is committing to a forecast you cannot make.
- Evaluation validity threat (track N5, evidence and evaluations) is a controlled list of named conditions — evaluation awareness, sandbagging, alignment faking, metagaming or grader-gaming, and reward hacking — under which behaviour during evaluation may not reflect behaviour in deployment. Metric choice, the mirage paper’s subject, is the measurement-side counterpart: a result can be invalid because the model behaved differently under test, or because the number you computed from its outputs did not mean what you took it to mean. A forecast inherits both. See the N5 evaluation validity threats guide for the full list.
Frequently Asked Questions
Are emergent abilities in large language models real?
It depends which claim you mean. Sharp jumps in reported benchmark scores are real and reproducible, but Schaeffer, Miranda and Koyejo (NeurIPS 2023) showed that for fixed model outputs the sharpness frequently comes from discontinuous metrics such as exact-string-match rather than from a discontinuity in model behaviour. Wu and Lo (ICLR 2025) offer a third reading in which the aggregate jump is genuine but arises from opposing difficulty-dependent trends cancelling out.
Does the “mirage” paper mean capability jumps are predictable?
No. It shows that apparent sharpness is often metric-dependent. Predictability is a separate question, and the same lead author’s 2024 follow-up, “Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?”, argues that downstream benchmark scores remain hard to forecast because the transformations from log-likelihood to score degrade the statistical relationship with scale.
How far ahead can frontier capability currently be forecast?
Short distances, with caveats. The GPT-4 Technical Report predicted final loss from models using at most 10,000× less compute and a HumanEval subset pass rate from at most 1,000× less compute, while stating that certain capabilities remain hard to predict. Snell and colleagues (November 2024) predicted emergence for models trained with up to 4× more compute, in some cases. Pimpale and colleagues (February 2025) forecast agent benchmark scores roughly a year out and flagged that their forecasts may be too conservative.
Why do safety frameworks use evaluation cadence instead of forecasts?
Because dense sampling substitutes for prediction. Google DeepMind’s Frontier Safety Framework commits to early warning evaluations run “frequently enough that we have notice before that threshold is reached,” and Anthropic’s RSP v3.0 (24 February 2026) extended its evaluation interval to six months to improve elicitation quality. Both are cadence decisions, not forecasting decisions. OpenAI’s Preparedness Framework v2 goes the other way and treats a forecast of Critical capability as itself sufficient to trigger development-time safeguards.
Why does the EU AI Act use a FLOP number rather than a capability test?
It uses both. Article 51(1) is a capability test; Article 51(2) adds a rebuttable presumption that a model has high-impact capabilities when cumulative training computation exceeds 1025 FLOP. Compute is observable in advance and not subject to metric disputes, which is exactly what a capability test lacks. Article 51(3) lets the Commission amend the thresholds by delegated act so they reflect the state of the art — an acknowledgement that the proxy drifts.
Does any of this apply to universities rather than frontier labs?
Partly. Compute thresholds in the EU AI Act and in export-control policy do not exempt academic training runs, so research computing units benefit from logging cumulative training FLOP during a run. Beyond that, the transferable lesson is structural: any institutional policy that gates a tool on a capability claim inherits the same forecasting weakness, and the practical mitigation is a fixed review cadence rather than a one-time assessment.
Primary Sources
- Wei et al., “Emergent Abilities of Large Language Models,” arXiv:2206.07682 (15 June 2022), published in Transactions on Machine Learning Research.
- Schaeffer, Miranda and Koyejo, “Are Emergent Abilities of Large Language Models a Mirage?”, arXiv:2304.15004 (28 April 2023); NeurIPS 2023, outstanding paper award.
- Schaeffer, Schoelkopf, Miranda, Mukobi, Madan, Ibrahim, Bradley, Biderman and Koyejo, “Why Has Predicting Downstream Capabilities of Frontier AI Models with Scale Remained Elusive?”, arXiv:2406.04391 (6 June 2024).
- Ganguli, Hernandez, Lovitt et al., “Predictability and Surprise in Large Generative Models,” arXiv:2202.07785; ACM FAccT 2022.
- Ruan, Maddison and Hashimoto, “Observational Scaling Laws and the Predictability of Language Model Performance,” arXiv:2405.10938 (May 2024).
- Snell, Wallace, Klein and Levine, “Predicting Emergent Capabilities by Finetuning,” arXiv:2411.16035 (25 November 2024).
- Wu and Lo, “U-shaped and Inverted-U Scaling behind Emergent Abilities of Large Language Models,” arXiv:2410.01692; ICLR 2025.
- Pimpale, Højmark, Scheurer and Hobbhahn, “Forecasting Frontier Language Model Agent Capabilities,” arXiv:2502.15850 (February 2025).
- OpenAI, “GPT-4 Technical Report,” arXiv:2303.08774, §“Predictable Scaling.”
- Google DeepMind, “Introducing the Frontier Safety Framework.”
- Anthropic, Responsible Scaling Policy update log; RSP v3.0 effective 24 February 2026 and RSP v3.4 effective 8 July 2026.
- OpenAI, Preparedness Framework version 2, 15 April 2025.
- METR, Common Elements of Frontier AI Safety Policies, 16 December 2025.
- Regulation (EU) 2024/1689 (AI Act), Article 51.








