Skip to main content
v2026.11,858 entries · CC-BY 4.0

External Review, Unpacked: Which Frontier Labs Let Outsiders Audit Their Safety Cases

Anthropic’s Long-Term Benefit Trust can compel outside review of its Risk Reports. California now requires labs to disclose how much third-party evaluation they used. Most other frontier labs’ safety policies still leave external review optional, or don’t mention it at all. This guide compares, lab by lab, who actually lets outsiders check their Responsible Scaling Policy compliance and who is still grading its own homework.

Written and maintained by CASRAI Editorial Board

Last updated

Short answer: almost no frontier AI lab currently submits its Responsible Scaling Policy (RSP) compliance to a binding, independent external audit. Anthropic is the one lab whose policy gives an outside body real, structural authority to compel and approve external review; several others describe third-party evaluation as something they may do, at their own discretion, without committing to it; and at least two labs’ public safety frameworks don’t specify an external-review mechanism at all. Last verified: September 20, 2026. CASRAI’s NIKOLAI dictionary names this specific gap — a record of “an assessment performed by a party outside the model developer” — as its own element, External Review, in the N8 Transparency and Review track; this guide is the first place on casrai.org to lay out, lab by lab, what that element would actually have to record.

CASRAI’s separate guide on third-party evaluator standards and methodology covers who evaluators are, how their independence is judged, and what access and publication rights they typically get once a lab brings them in. This guide asks the question one level up: which labs commit, in writing, to bringing an outside evaluator in for RSP-style compliance review at all — versus relying on their own internal teams to assess and attest to their own compliance. The two guides are meant to be read together, not as duplicates of each other.

Self-Attestation vs. External Review: What Actually Differs

“Self-attestation” means the lab that wrote the safety policy is also the body that checks and certifies its own compliance with that policy — typically an internal safety or policy team, reporting up an internal chain, with no outside party required to see the underlying evidence. “External review,” in the sense NIKOLAI’s element tracks, means a party outside the model developer — a person, an institution, or a government body — performs an assessment of a model, a risk report, a specific safeguard, or the developer’s compliance with its own framework, and that assessment’s type, scope, and output get recorded somewhere. The distinction that matters in practice is not “did a lab mention third parties,” almost every current framework does, but three sharper questions:

  • Compelled or discretionary? Can an outside party (or an internal-but-independent body) require a review to happen, or does the lab decide case by case whether to invite one in?
  • Unredacted or filtered access? Does the reviewer see the material a lab’s own team sees, or only a version the lab has already edited?
  • Binding or advisory? Does the review’s outcome constrain what the lab can do next, or is it feedback the lab is free to set aside?

Graded against those three questions, “external review” turns out to span a wide range — from a governance body with standing authority to demand it, down to a sentence in a policy document saying it might happen “when available and feasible.”

The Lab-by-Lab Comparison

The table below reflects each lab’s own published safety-framework language, as CASRAI could verify it directly or, where noted, as it is quoted and cited in NIKOLAI’s own crosswalk for the External Review element. It is not a ranking of which lab is “more safe” — it is a comparison of one specific structural feature: whether the lab’s own policy creates any outside role in checking its compliance.

Lab What the policy says about outside review Compelled or discretionary Source
Anthropic RSP v3.2 (April 2026) authorizes the Long-Term Benefit Trust (LTBT) — a body outside Anthropic’s normal corporate governance — to request external review of Risk Reports and to approve which external reviewers are used. RSP v3.4 (July 2026) clarifies that a Risk Report’s unredacted sections can be split across multiple external reviewers, so long as every part is seen by at least one of them. Compelled: the LTBT, not Anthropic’s own safety team, decides whether review happens and who does it. Anthropic Responsible Scaling Policy, versions 3.2 and 3.4 (anthropic.com/rsp), verified directly
OpenAI Per NIKOLAI’s crosswalk citation, OpenAI’s Preparedness Framework states it will work with third parties to evaluate models “when available and feasible,” and its Frontier Governance Framework says OpenAI “may solicit and obtain input from external experts.” Discretionary: both quotes are conditional (“when feasible,” “may”), with no outside body holding authority to compel review. OpenAI Preparedness Framework v2 §5.2 / Frontier Governance Framework §5, as cited in NIKOLAI’s External Review crosswalk (not independently re-fetched this session — OpenAI’s own site returned a 403 to this guide’s verification pass)
Google DeepMind DeepMind’s public safety materials confirm an active Frontier Safety Framework and, as of August 2026, a named pilot (“Piloting the world’s first double-blind AI evaluations”) that points toward more independent evaluation practice. NIKOLAI’s crosswalk separately cites the Framework’s own language on “external safety testing” by “specialist independent groups,” which this guide could not independently re-fetch. Discretionary as far as this guide could independently confirm; the double-blind pilot suggests movement, not yet a standing compelled-review commitment. deepmind.google/about/responsibility-safety, verified directly; framework-text quotes as cited in NIKOLAI’s crosswalk
G42 NIKOLAI’s crosswalk cites G42’s Frontier Safety Framework as committing to “annual external audits to verify compliance with the Framework.” Stated as compelled and recurring (annual) in the cited language — this guide could not independently re-fetch G42’s framework document to confirm the full context. G42 Frontier Safety Framework s.5, as cited in NIKOLAI’s External Review crosswalk
Meta and xAI NIKOLAI’s own External Review element lists both as “related, not mapped” — meaning CASRAI could not find language in either lab’s public framework specific enough to map cleanly against the element’s definition, one way or the other. Not established either way from public material, per NIKOLAI’s own assessment. NIKOLAI element page, casrai.org/nikolai/element/external-review, verified live

Reading the table honestly: Anthropic is the only row this guide could verify directly against the lab’s own primary document, and it is also the only row describing a body with standing authority to compel review, not just permission to allow it. Every other row either rests on NIKOLAI’s own citation of framework language this guide’s own verification pass could not re-confirm against the lab’s site directly (OpenAI, DeepMind’s specific quote, G42), or reflects the honest absence of a clear public commitment at all (Meta, xAI). That unevenness — one well-documented case, several partially-sourced ones, and two open questions — is itself the finding.

The Regulatory Backstops: California SB 53 and the FRONTIER Act

Two policy mechanisms outside any single lab’s own framework push toward more outside involvement, and it is worth being precise about what each actually requires, since it is easy to overstate both.

California SB 53 requires large frontier developers’ published safety frameworks to address “using third parties to assess the potential for catastrophic risks and the effectiveness of mitigations of catastrophic risks” (Cal. Gov’t Code §22757.12(a)(5), verified directly against the bill text), and requires transparency reports accompanying new or substantially modified frontier models to disclose “the extent to which third-party evaluators were involved” (§22757.12(c)(2)(C)). Read carefully, SB 53 is a disclosure mandate, not an external-audit mandate: a developer that used zero third-party evaluators is still complying, as long as it says so. What SB 53 forecloses is the option of staying silent about the question.

The FRONTIER Act (H.R. 9925, 119th Congress) is the one proposal that would go further, requiring “very large frontier developers” to grant an Independent Verification Organization access to unredacted materials not less than once every six months. It has not been enacted. CASRAI’s own comparison, Mandatory vs. Voluntary: Who Has to Let Outsiders Look, works through the FRONTIER Act’s mechanics against voluntary lab policies in full; this guide’s table above is the lab-by-lab layer that comparison doesn’t itself provide.

Who Actually Does This Work: METR and the Evaluator Ecosystem

A framework saying a lab “may” use third parties only means something if a credible third party exists to do the work. METR is the organization most frequently named in this role: its own site describes its work as including “independent reviews of AI developers’ risk assessments” alongside its own capability evaluations, verified directly against metr.org as of this guide’s publication. CASRAI’s separate guide, What Is METR?, covers its origin, funding, and methodology in depth.

It’s worth naming a limit here too, rather than assuming the evaluator ecosystem is further along than it is. The Frontier Model Forum — an industry body Anthropic, Google DeepMind, Microsoft, and OpenAI formed — published a technical report specifically on frontier capability assessments that says, in its own words, it “does not cover methods for assessing the effectiveness of new safeguards against these risks, determining alignment of models, monitoring for harm post-deployment, or conducting third-party assessments,” and treats “external domain experts” as consultants brought in to supplement a developer’s own internal assessment team, not as independent verifiers of it. That is a direct, verified statement from an industry self-governance body that standardized third-party assessment methodology does not yet exist even in its own published guidance — a useful check against assuming the infrastructure for external review is more mature, industry-wide, than the lab-by-lab table above actually shows.

Where NIKOLAI Fits: CASRAI’s External Review Element

Everything compared above is what NIKOLAI’s External Review element, in its N8 (Transparency and Review) track, is built to record as a structured fact rather than a scattered policy sentence: “a record of an assessment performed by a party outside the model developer — of a model, a risk report, a safeguard, or compliance with a framework — capturing the review’s type, scope, and output,” per the element’s definition as verified live on its own page. As of this guide’s publication the element carries status Proposed in NIKOLAI’s current release (nikolai-v0.2).

It is worth stating plainly, because it’s easy to blur: NIKOLAI is CASRAI’s own project — an independent, unendorsed reference mapping, not a standard any of the labs above have reviewed or agreed to. Every row in NIKOLAI’s own crosswalk table for this element, including the Anthropic row, is what NIKOLAI itself calls a shadow mapping: CASRAI’s own editorial reading of publicly available material, carrying confidence labels (exactEQ, closeCL) that reflect how much interpretation went into each one, not a mapping the named organization has confirmed. None of Anthropic, OpenAI, Google DeepMind, the EU, California, NIST/AISI, METR, the Frontier Model Forum, or G42 — all eleven rows currently on NIKOLAI’s live crosswalk — has filed a Mapping Declaration endorsing how NIKOLAI classifies their language. If one does, NIKOLAI updates that row from shadow mapping to declared, and this guide will note the change.

Before the External Review element existed as a distinct entry, this specific piece of NIKOLAI’s N8 track had no dedicated treatment anywhere on casrai.org — the closest existing coverage was a single sentence, in the third-party-evaluator guide, quoting the track’s own rationale that external review “only means something if the evaluator is independent and a developer can’t quietly redact an inconvenient finding.” That sentence is still true and still the right one-line summary of why the element exists; this guide is the fuller treatment it pointed toward but didn’t provide.

Frequently Asked Questions

What counts as “external review” of RSP compliance, as opposed to self-attestation?

External review means a party outside the model developer — a person, an institution, or a government body — independently assesses a model, a risk report, a specific safeguard, or the developer’s compliance with its own safety framework. Self-attestation means the developer’s own internal team performs that same assessment and certifies the result itself, with no outside party required to see the underlying evidence. Most current frontier-lab frameworks sit closer to self-attestation, with external review available as an option rather than a requirement.

Does any frontier lab currently undergo a mandatory, binding external audit of its safety-policy compliance?

Not in the sense of a government-enforced, statutory audit — no such requirement is currently in force in the US or EU as of this guide’s publication. Anthropic’s policy comes closest among voluntary lab commitments: its Long-Term Benefit Trust has standing authority under the RSP to request external review of Risk Reports and approve the reviewers used, which is a real internal governance mechanism, not a government mandate.

Does California’s SB 53 require frontier labs to use external evaluators?

No. SB 53 requires large frontier developers’ published safety frameworks to address their approach to third-party risk assessment, and requires transparency reports to disclose how much third-party involvement a given model’s catastrophic-risk assessment actually had. A developer that used no third-party evaluators is still complying with SB 53, as long as it discloses that. SB 53 mandates transparency about the answer, not a particular answer.

Is CASRAI’s NIKOLAI project an official or endorsed framework for classifying external review?

No. NIKOLAI is CASRAI’s own, independent, unendorsed reference dictionary. Its External Review element and crosswalk table represent CASRAI’s own editorial reading of publicly available lab and regulatory material, not a mapping any named organization has reviewed, confirmed, or endorsed. Every crosswalk row is a shadow mapping unless a Mapping Declaration says otherwise.

Where does METR fit into external review of RSP compliance?

METR is one of the most frequently named independent organizations performing this kind of work, describing its own role as including independent reviews of AI developers’ risk assessments alongside its capability evaluations. It is an evaluator labs can choose to engage, not a body with standing authority to compel review at any lab currently covered in this guide.

Related Reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about External Review, Unpacked: Which Frontier Labs Let Outsiders Audit Their Safety Cases

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →