Skip to main content
v2026.11,858 entries · CC-BY 4.0

Direct comparison

EFF vs. AI Now on AI Safety Frameworks

EFF and the AI Now Institute both critique AI lab safety practices, but for opposite reasons. Compare their arguments, evidence, and what each would change.

Written and maintained by CASRAI Editorial Board

Last updated

Ask CASRAI · free to try

Ask about EFF vs. AI Now on AI Safety Frameworks

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

How do EFF, AI Now Institute compare side by side?

The table below compares EFF, AI Now Institute across 6 procurement-relevant dimensions, from core diagnosis through where the two implicitly agree.

Side-by-side comparison

DimensionEFFAI Now Institute
Core diagnosisThe gap is a failure to apply cybersecurity practice that already exists, not an absence of AI-specific rules. Jacob Hoffman-Andrews, Tori Noble, and Maddie Daly wrote (Sept. 17, 2026) that following “fundamental best practices would have prevented or substantially mitigated all of the incidents at AI labs that we currently know about.”The gap is that labs are moving away from rigorous methodology that already existed. Heidy Khlaaf's report (April 21, 2025) states: “We're seeing the erosion of tried-and-true evaluation approaches in favor of vague claims of capabilities that fail to meet even the most basic safety thresholds.”
What “the existing standard” meansGeneral information-security discipline: sandboxing, monitoring, logging, and vulnerability handling — practices that apply to any software system, not something unique to AI.Quantifiable, externally legible risk-threshold methodology of the kind historically used for nuclear and other high-consequence engineering systems, not self-graded qualitative judgment calls.
View on writing new, AI-specific rulesSkeptical. “Minimum safety requirements specific only to current AI technologies are likely to become obsolete; legal standards tied to well-established cybersecurity best practices are far more likely to stand the test of time.” Wants lawmakers to extend existing cybersecurity law, not create a bespoke AI carve-out.Does not argue for or against a specific new statute. The target is the substitution happening inside labs' own self-published frameworks — the erosion is the problem, regardless of which body is supposed to be setting the bar.
How it lands on RSP / Preparedness Framework specificallyDoesn't name either framework directly. Its logic implies that if labs were actually following standard InfoSec practice, voluntary commitments like RSP's model-weight-security provisions should already cover much of what regulators need — baseline cybersecurity law is the priority, not auditing the frameworks' internal design.Names the pattern directly: a framework that lets a lab define its own capability threshold qualitatively, then assess its own model against that self-defined bar, is the exact “vague claims of capabilities” mechanism her report identifies. That structure is common to both RSP and the Preparedness Framework — see CASRAI's RSP vs. Preparedness Framework vs. FSF comparison for how each currently defines a threshold.
Stakes, as framed by each sourceRegulatory durability — avoiding legislation that ages out as fast as the underlying technology changes.National security. Khlaaf's title frames weakened thresholds as a “self-fulfilling prophecy,” arguing the erosion itself increases the risk it was supposed to manage, particularly around accelerated military AI deployment without adequate testing.
Where the two implicitly agreeThat today's voluntary, lab-authored safety commitments are not sufficient on their own, as demonstrated by incidents that a stricter baseline would have caught.That today's voluntary, lab-authored safety commitments are not sufficient on their own — though because they've drifted from a higher standard that was already available, not because no standard exists.

Common questions

Common questions about EFF vs AI Now Institute

Do EFF and the AI Now Institute actually disagree, or are they making the same argument two ways?

+

They share a conclusion — current practice at AI labs is inadequate — but reach it from opposite premises. EFF's argument is that a sufficient standard already exists in general cybersecurity practice and just isn't being applied or legislated; the fix is extending that existing standard, not writing AI-specific rules. AI Now's argument is that a more rigorous, AI-relevant standard already existed (in the risk-threshold-setting tradition it borrows from) and labs have moved away from it toward vaguer, self-graded language. One calls for applying an existing general standard; the other calls for restoring an existing specialized one.

Does either critique target Anthropic's RSP or OpenAI's Preparedness Framework by name?

+

Neither piece names RSP or the Preparedness Framework directly. AI Now Institute's critique of “vague claims of capabilities” that “fail to meet even the most basic safety thresholds” describes the qualitative, lab-defined threshold structure that both frameworks use, as documented on CASRAI's RSP vs. Preparedness Framework vs. FSF comparison. EFF's piece is framed around cybersecurity legislation generally and doesn't address frontier-lab scaling policies as a category.

Does CASRAI take a position on which critique is correct?

+

No. CASRAI's own NIKOLAI project, an independent, unendorsed reference dictionary of frontier-AI-safety elements, documents how named organizations and jurisdictions define and disclose their capability thresholds under its N3 track element, Capability Threshold — as a factual record readers can use to evaluate either argument for themselves. It does not adjudicate whether EFF's cybersecurity-first framing or AI Now's threshold-erosion framing is correct, and every crosswalk row it publishes is CASRAI's own interpretive reading, not a mapping any named organization has agreed to.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →