Skip to main content
v2026.11,858 entries · CC-BY 4.0

Is AI Safety a Property of the Model? The Narayanan-Kapoor Critique

Arvind Narayanan and Sayash Kapoor’s March 2024 essay ‘AI Safety Is Not a Model Property’ argues that capability-threshold frameworks — Anthropic’s RSP, OpenAI’s Preparedness Framework, Google DeepMind’s FSF — rest on a category error: safety depends on deployment context, not a property a model can be tested for and certified to have. This guide walks through their argument, in full attribution, as the academic counterpoint this cluster’s framework-explainer and comparison pages have been missing.

Written and maintained by CASRAI Editorial Board

Last updated

Last verified: September 20, 2026. Every framework this cluster has covered so far — Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework, Google DeepMind’s Frontier Safety Framework, and the structural comparison across all three — shares one underlying premise: that a model’s dangerous capability can be measured, that crossing a defined threshold is the meaningful trigger for stronger safeguards, and that the lab publishing the framework is the right party to run the tests that decide when the threshold has been crossed. Princeton computer scientists Arvind Narayanan and Sayash Kapoor published a direct challenge to that premise on March 12, 2024, in an essay titled “AI Safety Is Not a Model Property.” The essay first appeared on their newsletter AI Snake Oil; the newsletter has since rebranded as Normal Technology and now hosts the piece at normaltech.ai. Nothing about the argument itself has changed in that move — only the address it lives at. CASRAI’s own NIKOLAI dictionary has an element built around exactly the who-runs-the-test question this essay raises — more on that below.

Narayanan and Kapoor are not frontier-lab insiders publishing a rival framework. They are academic researchers — Narayanan a Princeton computer science professor, Kapoor a Princeton PhD researcher, co-authors of the AI Snake Oil book and newsletter — writing as outside critics of the entire model-centric safety paradigm this cluster has otherwise documented from the inside: what the labs themselves say their frameworks do. This guide is CASRAI’s attempt at the missing companion piece: what a credentialed academic critique of that paradigm actually argues, in full attribution, without CASRAI editorializing the argument into something stronger or weaker than Narayanan and Kapoor themselves wrote.

The Core Claim: Safety Depends on Context, Not the Model

The essay’s title is also its thesis. Narayanan and Kapoor write plainly: “Safety depends to a large extent on the context and the environment in which the AI model or AI system is deployed.” Their argument is that treating safety as a property intrinsic to a model — something a lab can test for in isolation and certify before release, the way a capability-threshold framework structurally assumes — misdescribes what actually determines whether an AI system causes harm. A model that behaves safely in a lab’s own evaluation harness can still be dangerous once deployed into a context the evaluation didn’t anticipate, and a model that looks risky in isolated testing may be harmless in a deployment context with strong surrounding controls. As they put it, “even if models themselves can somehow be made ‘safe’, they can easily be used for malicious purposes” — the model-level property, even where achievable, doesn’t settle the question a framework needs it to settle.

That is a direct challenge to the structural logic this cluster’s capability-threshold guide documents across fourteen organizations: a framework that defines a threshold, tests the model against it, and applies safeguards once the model crosses it is, in Narayanan and Kapoor’s framing, asking the model to carry a property — safety — that belongs to the deployment context instead.

What It Means for RSPs, Preparedness Frameworks, and the FSF

Narayanan and Kapoor single out capability-threshold mechanics specifically, not just the general concept of AI safety testing. On the compute-based thresholds that show up across governance proposals — most visibly the 1026 FLOP training-compute trigger used in the since-rescinded US executive order and echoed in several state and international proposals — they argue policymakers “seem to have converged on 1026 rather arbitrarily,” without a demonstrated causal link between that specific compute figure and the harms the threshold is meant to catch. Their target there is a compute-based threshold specifically, a mechanism distinct from the capability-based thresholds Anthropic, OpenAI, and Google DeepMind each define in their own frameworks — but the underlying critique reaches further: if a number this consequential was picked without a rigorous basis, that is evidence for their broader claim that thresholds get treated as more scientifically grounded than the process that produced them actually supports.

Applied to the three frameworks this cluster has already covered, the critique lands somewhat differently on each. Anthropic’s RSP thresholds are capability-based rather than compute-based — for example, the AI R&D threshold’s 5x-cost-substitution test — which sidesteps the specific 1026 arbitrariness charge Narayanan and Kapoor level at compute triggers. But it does not sidestep their larger point: a lab-run evaluation of whether a model can substitute for a research engineer still tests a property of the model in a controlled setting, not whether the model causes harm once deployed into the actual mix of user population, surrounding tooling, and misuse countermeasures that determine real-world outcomes. The same applies to OpenAI’s Preparedness Framework and Google DeepMind’s FSF: both define their own capability thresholds and test against them internally, which is precisely the model-property structure Narayanan and Kapoor’s essay argues is the wrong shape of intervention.

Where They Argue Defenses Actually Belong

If safety isn’t a model property, the essay’s constructive claim is about where effort should go instead: outside the model. Narayanan and Kapoor point to mechanisms like email scanners and URL blacklists as examples of the kind of deployment-context defense they mean — controls that operate on how a system is used, not on what the underlying model can theoretically do. Their broader claim is that retrospective detection of misuse is, in their assessment, technically easier to build and verify than model-level alignment is — catching and responding to harmful use after the fact, in the deployment environment, rather than trying to make the model itself incapable of ever producing a harmful output in principle.

That is a genuinely different emphasis from what a capability-threshold framework is built to do. An RSP, a Preparedness Framework, and the FSF are all pre-deployment gates: they test the model, and they hold if the model fails the test. Narayanan and Kapoor’s argument, if taken seriously, doesn’t say those gates are worthless — it says they are necessarily incomplete on their own, because the harm a model-property test is trying to prevent is determined jointly by the model and by everything sitting around it once deployed, and only the deployment-side controls can respond to the parts of that joint determination a pre-release model test structurally cannot see.

Who Should Run the Red Team

The essay’s third strand is about who should be doing the testing, not just what should be tested. Narayanan and Kapoor argue that “red teaming should be led by third parties with aligned incentives” rather than by the lab whose own framework and business incentives are on the line — and that red-teaming work should focus on “advancing frontier of adversary capabilities” generally, rather than narrowly on the safety of one lab’s own model. They also flag a funding-concentration concern specific to this field: work in effective-altruism-adjacent biosecurity risk research clusters around a small set of funders, which they raise as a reason to scrutinize whether red-teaming incentives are actually independent of the interests — commercial or philanthropic — of whoever is paying for the evaluation.

This is the strand of the essay with the most direct tie to material CASRAI has already published about this cluster’s frontier-safety landscape, covered next.

Where NIKOLAI Fits In

NIKOLAI, CASRAI’s own independent, unendorsed reference dictionary for frontier AI safety, has an element built around exactly the question Narayanan and Kapoor raise about who runs a red team: Evaluator Independence and Conflict of Interest, part of Track N8, Transparency and Review. NIKOLAI’s own definition — verified directly against the live element page on September 20, 2026 — describes a record of “declared financial, organisational and personal relationships between an evaluator and the developer being evaluated, and the independence test applied to clear the evaluator for the engagement,” and the page is explicit that this is “CASRAI’s own proposed definition, not a definition any named organisation has agreed to.”

Narayanan and Kapoor’s essay isn’t a framework document with its own named independence test the way Anthropic’s RSP or METR’s conflict-of-interest policy are — it’s an academic argument for why third-party-led red-teaming with aligned incentives matters, not a specification NIKOLAI’s element could crosswalk as a source row the way it crosswalks Anthropic, the EU, METR, the Frontier Model Forum, and two Congressional bills in CASRAI’s own evaluator-independence guide. So this page draws the connection narratively rather than adding a formal shadow-mapping row: Narayanan and Kapoor supply the academic argument for why evaluator independence matters at all — misaligned incentives, including funding concentration, distort what a red team is willing to find and report — and NIKOLAI’s N8 element, plus the seven organizations and bills already mapped against it, is CASRAI’s attempt to make that independence checkable in practice, one published policy at a time. Reading the two together is the point: the essay explains why the question matters; the evaluator-independence guide is where CASRAI has actually gone and checked who answers it, and how.

Frequently Asked Questions

Who are Arvind Narayanan and Sayash Kapoor?

Arvind Narayanan is a computer science professor at Princeton University; Sayash Kapoor is a Princeton PhD researcher. Together they write AI Snake Oil (now rebranded Normal Technology), a newsletter and book on AI hype and AI safety, and are academic critics of the frontier-AI-safety field rather than employees of any lab whose framework they critique.

What does “AI Safety Is Not a Model Property” actually argue?

Published March 12, 2024, the essay argues that safety depends on the deployment context an AI system is used in, not on a property intrinsic to the model that a lab can test for and certify before release. It argues defenses should focus more on the deployment environment — retrospective misuse detection, for example — and that red-teaming should be led by independent third parties with aligned incentives rather than by the lab whose own model and framework are being assessed.

Does this critique apply equally to Anthropic’s RSP, OpenAI’s Preparedness Framework, and Google DeepMind’s FSF?

The essay’s specific arbitrariness charge targets the 1026 FLOP compute threshold used in some governance proposals, which is a different mechanism from the capability-based thresholds Anthropic, OpenAI, and Google DeepMind each define in their own frameworks. But the broader argument — that a pre-deployment, lab-run test of model capability cannot by itself determine real-world safety, because safety is jointly determined by the model and its deployment context — applies to the model-property structure all three frameworks share, not to compute thresholds specifically.

Is NIKOLAI’s evaluator-independence element based on the Narayanan-Kapoor essay?

No. NIKOLAI’s N8 element, Evaluator Independence and Conflict of Interest, is CASRAI’s own independent editorial synthesis, crosswalked against published policies from Anthropic, the EU, METR, the Frontier Model Forum, and two Congressional bills — not against the Narayanan-Kapoor essay, which is not a policy document and isn’t one of the element’s crosswalk rows. The connection this guide draws is thematic: the essay argues why third-party-led evaluation with aligned incentives matters; NIKOLAI’s element is CASRAI’s attempt to track which organizations have published an actual independence test.

Related Reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about Is AI Safety a Property of the Model? The Narayanan-Kapoor Critique

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →