Skip to main content
v2026.11,858 entries · CC-BY 4.0

Direct comparison

Bengio-Hinton vs. Narayanan-Kapoor on AI Risk

Bengio, Hinton and 23 co-authors call for binding AI governance; Narayanan and Kapoor say safety is contextual, not a model property. Compared with citations.

Written and maintained by CASRAI Editorial Board

Last updated

Ask CASRAI · free to try

Ask about Bengio-Hinton vs. Narayanan-Kapoor on AI Risk

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

How do Bengio, Hinton et al. (Science, 2024), Narayanan & Kapoor (2024) compare side by side?

The table below compares Bengio, Hinton et al. (Science, 2024), Narayanan & Kapoor (2024) across 6 procurement-relevant dimensions, from core diagnosis of the current gap through what each piece argues against.

Side-by-side comparison

DimensionBengio, Hinton et al. (Science, 2024)Narayanan & Kapoor (2024)
Core diagnosis of the current gap"Present governance initiatives lack the mechanisms and institutions to prevent misuse and recklessness, and barely address autonomous systems." The paper frames the problem as an absence of binding regulatory authority, not a conceptual error in how labs test their models."Safety depends to a large extent on the context and the environment in which the AI model or AI system is deployed." Their diagnosis is structural: treating safety as something a lab can test for and certify in a model before release misdescribes what actually determines harm.
What capability thresholds are forEndorses capability-triggered rules in principle, provided they adapt: "The key is policies that automatically trigger when AI hits certain capability milestones. If AI advances rapidly, strict requirements automatically take effect, but if progress slows, the requirements relax accordingly."Directly criticizes the compute-based version of this mechanism. On the 10^26 FLOP training-compute trigger used in several state and international proposals, they argue policymakers "seem to have converged on 10^26 rather arbitrarily," without a demonstrated causal link to the harms the threshold is meant to catch.
Where enforcement authority should sitWith government. "Governments must be prepared to license their development, restrict their autonomy in key societal roles, halt their development and deployment in response to worrying capabilities, mandate access controls, and require information security measures robust to state-level hackers." Regulators should also "hold frontier AI developers and owners legally accountable for harms from their models that can be reasonably foreseen and prevented."With independent evaluators, not government licensing power. Their essay argues "red teaming should be led by third parties with aligned incentives" rather than by the lab whose own framework and business interests are on the line -- a narrower, evaluation-focused prescription than Bengio et al.'s state-licensing regime.
Role of external audits and accessCalls for regulators to require that "frontier AI developers grant external auditors on-site, comprehensive ('white-box'), and fine-tuning access from the start of model development," alongside registration of frontier systems and mandatory incident reporting and whistleblower protections.Proposes no specific audit-access regime. Their closest analogue is the red-teaming independence argument above, plus a caution that red-teaming funding itself clusters around a small set of funders -- a reason, they argue, to scrutinize whether evaluation incentives are genuinely independent.
Where the actual defense should be builtAcross the whole lifecycle -- upstream restrictions on development plus downstream monitoring, incident reporting, and access controls -- rather than any single point of intervention.Primarily in the deployment context, after release. They point to mechanisms like email scanners and URL blacklists as the kind of control that operates on how a system is used, arguing retrospective misuse detection is, in their assessment, technically easier to build and verify than model-level alignment.
What each piece argues against"Present governance initiatives" broadly -- the absence of binding rules across jurisdictions, not any one lab framework by name.The structural logic shared by Anthropic's RSP, OpenAI's Preparedness Framework, and Google DeepMind's FSF specifically: each defines its own capability threshold and tests against it internally, which Narayanan and Kapoor argue is the wrong shape of intervention regardless of where the threshold is set.

Common questions

Common questions about Bengio, Hinton et al. (Science, 2024) vs Narayanan & Kapoor (2024)

Are Bengio, Hinton and co-authors responding directly to Narayanan and Kapoor, or vice versa?

+

No direct evidence of either. The Science paper was first submitted October 26, 2023 and last revised May 22, 2024; Narayanan and Kapoor's essay published March 12, 2024, landing between those two dates. Neither piece cites the other by name, and CASRAI found no reference from either side to the other's argument in the published text. They are contemporaneous, independently argued positions, not a direct exchange.

Does Narayanan and Kapoor's arbitrariness critique of capability thresholds also apply to Bengio et al.'s proposal?

+

Only partially. Narayanan and Kapoor's specific target is the 10^26 FLOP compute-based trigger used in some governance proposals. Bengio et al. don't defend that specific figure -- their proposal is for rules tied to capability milestones generally, adapting automatically as progress speeds up or slows down. That adaptiveness answers a different worry (rules going stale) than the one Narayanan and Kapoor raise (a threshold number picked without a demonstrated causal basis).

Is one of these the 'establishment' position and the other the 'fringe' one?

+

No -- treating it that way misreads both. Bengio, Hinton and their 23 co-authors are a high-profile group, including two Turing Award winners and a Nobel laureate economist, publishing in a peer-reviewed general-science journal -- but the paper is a position piece with named signatories, not an enacted policy or a body with regulatory authority. Narayanan and Kapoor are credentialed academic computer scientists at Princeton, not industry insiders or outside amateurs, offering a structural critique of the model-centric paradigm the Science paper's own recommendations largely assume. Neither position has been adopted as binding law in the form either group describes.

Does either paper reference NIKOLAI, or has NIKOLAI adopted either position?

+

No. Neither the Bengio-Hinton paper (Science, 2024) nor the Narayanan-Kapoor essay mentions NIKOLAI, which did not exist when either was written. CASRAI's own NIKOLAI project, an independent, unendorsed reference dictionary of frontier-AI-safety elements, has a track built around the kind of adaptive-governance mechanism Bengio et al. call for: N9, Commitments and Governance, includes Framework Update Protocol -- 'a proposed rule set for when and how a safety framework may be revised, covering review cadence, triggers, approvers, and publication deadlines' -- verified directly against the live element page on September 20, 2026. N9 also includes Accountable Decision-Maker and Sign-Off, 'a proposed record of the named role or person who approves a risk, deployment, or framework decision, and what was approved, by whom, and when,' which echoes Bengio et al.'s call to hold a named party legally accountable. Both definitions are CASRAI's own editorial synthesis, crosswalked against Anthropic's RSP, the EU GPAI Code of Practice, and California SB 53 -- not a mapping either Bengio et al. or Narayanan and Kapoor themselves drew, endorsed, or were consulted on. It is offered here as a shadow mapping only: a CASRAI-authored lens for reading their disagreement, not a claim about what either paper's authors intended.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →