Written and maintained by CASRAI Editorial Board
Last updated
Every major frontier-AI safety framework acknowledges that published risk reports will sometimes have parts withheld — for trade secrets, national security, or plain competitive advantage. What they do not agree on is whether the developer has to tell anyone a redaction happened, or why. Built from NIKOLAI’s redaction crosswalk (Track N8, nine rows sourced from primary framework and statute text), this guide walks through what each major framework actually commits to on that question — and where the commitment simply stops.
The short version: Anthropic, Meta, California’s SB 53, and the independent evaluator METR all build a specific, named commitment into their process — disclose that a redaction happened, and say why, at least at a high level. The EU’s GPAI Code of Practice, and what NIKOLAI can currently verify about the UK’s approach, point the other way: information flows to the regulator largely unredacted, but no framework text found so far obligates a developer to tell the public, specifically, which parts of a report were cut and why.
What “redaction” means here
A redaction, in this context, is not silence. A framework that never mentions a risk area at all hasn’t redacted anything — it simply hasn’t addressed it. A redaction is a recorded removal: information that was going to be in the report, was pulled before publication, and (in the frameworks discussed below) is supposed to leave a visible mark where it used to be — a flagged section, a footnote, a line noting that material was omitted. NIKOLAI’s own working definition frames it as needing four things: where the removal happened, roughly how much was removed, why (drawn from a defined reason vocabulary — IP protection, national security, personal privacy), and who decided. That four-part framing is CASRAI’s own editorial proposal, not any organization’s declared position; the actual frameworks below vary in how many of those elements they require in practice.
The disclose-and-justify camp: Anthropic, Meta, California, METR
Anthropic’s Responsible Scaling Policy, version 3.4 (effective July 8, 2026) states the commitment directly, in its section on Risk Report publication: “We will disclose the existence of each redaction made in the public version of the report and aim to give a brief justification for such redactions, though in many cases this will necessarily be very high-level.” The RSP lists four categories a redaction can fall under — legal compliance, intellectual property protection, public safety, and privacy — and the policy’s July 2026 changelog entry (v3.4, change 4) makes clear the disclosure duty became an explicit, versioned requirement in this revision, noting that Anthropic’s previous Risk Report already met the new standard.
Meta’s Advanced AI Scaling Framework, version 2, uses almost identical language in its section on Preparedness Reports: “Preparedness reports will provide as much detail as possible, while redacting information to protect trade secrets or as appropriate under law. If we redact information, we will describe our reasons why to the extent that we can without creating undue risk.” Like Anthropic’s commitment, this covers both the fact of a redaction and a reason for it, with an explicit escape valve (“to the extent that we can”) if giving the reason would itself be risky.
California’s SB 53 — the Transparency in Frontier Artificial Intelligence Act — goes further than either lab framework, because it is a statute rather than a voluntary policy. Section 22757.12(f)(2) reads: “the frontier developer shall describe the character and justification of the redaction in any published version of the document to the extent permitted by the concerns that justify redaction and shall retain the unredacted information for five years.” That is a disclosure duty with two extra teeth a voluntary framework doesn’t carry: it is a legal “shall,” not a “we will aim to,” and it comes bundled with a document-retention requirement — a covered developer has to keep the unredacted version on file for five years, presumably so a regulator or plaintiff can eventually reach it even though the public never sees it. CASRAI’s separate guide on SB 53’s full transparency-report requirements covers the rest of the statute.
METR, the third-party evaluator, applies a version of the same norm to the incident reports it publishes about the labs it evaluates — effectively taking on the disclosure duty on the lab’s behalf. In its August 26, 2026 writeup investigating an OpenAI/Hugging Face incident, METR opens with a redaction summary statement: “Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions.” The mechanism is slightly different — METR is vouching for the completeness of a report it didn’t write, rather than disclosing its own redactions — but it serves the same function for a reader: nobody has to wonder whether an unstated redaction is hiding something material to the conclusions.
OpenAI: permitted, not promised
OpenAI’s Preparedness Framework, version 2 (last updated April 15, 2025) takes a visibly lighter stance. In its section on transparency and external participation, it states: “Such disclosures about results and safeguards may be redacted or summarized where necessary, such as to protect intellectual property or safety.” That sentence establishes that redaction is permitted and gives two example reasons — but unlike Anthropic’s or Meta’s text, it does not commit OpenAI to disclosing that a redaction occurred, or to explaining why, the way RSP §3.5 or AASF §2.2.1 do. The Preparedness Framework’s public-disclosures commitment focuses on describing the scope of testing performed, capability evaluations for each Tracked Category, and the reasoning for a deployment decision; redaction gets a permission in the same paragraph, not a matching disclosure duty.
The EU and UK: unredacted to the regulator, not necessarily to the public
The EU’s GPAI Code of Practice (Safety and Security Chapter) takes a structurally different approach: instead of a public disclose-and-justify duty, it routes unredacted material to the regulator and leaves public disclosure conditional.
- Framework notifications (Measure 1.4). “Signatories will provide the AI Office with (unredacted) access to their Framework, and updates thereof, within five business days of either being confirmed.” No national-security carve-out appears in this Measure.
- Model Report notifications (Measure 7.7). Signatories provide the AI Office with access to the Model Report “without redactions, unless they are required by national security laws to which Signatories are subject,” both at initial market placement and for any later update. Here the carve-out is explicit — but it runs toward the regulator, not the public: even the redacted case means the AI Office gets less, not that the public is told what was cut.
- Serious-incident reports (Measure 9.2). Information Signatories must report to the AI Office and national authorities is provided “to the best of their knowledge, redacted to the extent necessary to comply with other Union law applicable to such information” — a narrower, different carve-out (Union-law compliance, not IP or national security) applied to a different channel.
- Public transparency (Measure 10.2). This is the closest the Code comes to a public-facing rule, and it is conditional. Signatories publish a summarised Framework and Model Report(s) only “if and insofar as necessary to assess and/or mitigate systemic risks,” and that summary itself comes “with removals to not undermine the effectiveness of safety and/or security mitigations and to protect sensitive commercial information.” Nowhere in that Measure, on the text reviewed for this guide, is there language requiring a signatory to disclose the existence of a specific redaction, or justify it, to the public. The obligation is to publish a summary with permitted removals, not to flag each cut.
The pattern across all four Measures is consistent: the Code is built around getting complete, or near-complete, information to the AI Office, with narrow, purpose-specific carve-outs for national security and other Union law. It is not built around a public disclose-and-justify duty of the kind Anthropic’s RSP, Meta’s AASF, or California’s SB 53 spell out. That is a real, textual difference in what each document requires — not a claim that the EU permits more secrecy overall. If anything, the EU Code’s regulator-facing default is stricter than any lab framework’s public-facing one; it simply doesn’t extend that strictness to a public disclosure duty for individual redactions.
NIKOLAI’s crosswalk also carries a row for the UK AI Security Institute (AISI), but CASRAI is flagging it here, not asserting it: the only citation NIKOLAI’s researchers found for a UK AISI redaction practice is a secondhand mention — that UK AISI’s methodology was published “with implementation details redacted” — inside an Anthropic Risk Report, not a UK AISI primary document. That is an unverified pointer at the discovery stage, not a confirmed UK AISI rule, and it is presented on those terms.
xAI: a possible inconsistency, unconfirmed
NIKOLAI’s crosswalk flags one more data point, and it is explicitly provisional. xAI’s Risk Management Framework of August 20, 2025 reportedly described certain transparency items as ones xAI may publish in redacted form, with unredacted versions made available to vetted red teams or government agencies — a version-permits-redaction model roughly in the OpenAI mold, on NIKOLAI’s reading of that document. But NIKOLAI’s researchers found no redaction clause at all in xAI’s Frontier AI Framework dated June 30, 2026. That absence is worth flagging — it would mean xAI’s most recent framework dropped a commitment its predecessor had — but it is not being stated here as settled fact: the June 2026 document’s own PDF metadata reportedly reads “Privileged/Confidential DRAFT working FRAMEWORK DOC,” and no xAI statement was found disambiguating whether that document is a draft or xAI’s current, final framework. Until that status question is resolved, treat “xAI dropped its redaction clause between August 2025 and June 2026” as a provisional finding, not a confirmed one.
One more open thread: the AI Evaluator Forum
A discovery-level pass on the broader evaluator ecosystem turned up a group calling itself the AI Evaluator Forum (AEF-1), whose materials list “redactions” and “redaction disclaimer” as named sub-elements of something. NIKOLAI’s researchers have not read past the snippet naming them, so this is noted here purely as an open thread for a future pass — not as a claim about what AEF-1 actually requires.
The real divide
Strip away the framework-by-framework detail and one pattern holds: Anthropic, Meta, California’s SB 53, and METR all require someone outside the organization to be told, specifically, that a redaction happened and roughly why. Three of the four are public-facing commitments; SB 53 backs its version with a legal retention duty on top. The EU’s GPAI Code, and what can currently be verified of the UK’s approach, point the opposite way: they are built to get complete information to a regulator, with narrow carve-outs for national security or other law, but they don’t currently carry an equivalent public-disclosure mandate for the fact or reasoning behind a specific redaction. That isn’t a claim that the EU or UK permit more actual secrecy from regulators — the regulator-facing duty in the GPAI Code is arguably tighter than the public-facing one in any lab’s own framework. It’s a claim about who gets told what, and it is the sharpest single divergence NIKOLAI’s redaction crosswalk turns up.
NIKOLAI’s Redaction element
This guide is built directly from NIKOLAI’s Redaction element, part of Track N8 — Transparency and Review. NIKOLAI is CASRAI’s own frontier-AI-safety dictionary (current release nikolai-v0.2, 64 elements across 10 tracks), and it is worth repeating what that does and doesn’t mean here: every row in the Redaction element’s crosswalk — Anthropic’s, Meta’s, OpenAI’s, California’s, the EU’s, METR’s, xAI’s, and the still-unverified UK AISI and AI Evaluator Forum pointers — is what NIKOLAI calls a shadow mapping: CASRAI’s own independent reading of a document each organization has actually published, not a mapping any of those organizations has reviewed, endorsed, or declared. None of them has filed a Mapping Declaration confirming how it actually uses the term “redaction” internally. NIKOLAI is an independent reference work; CASRAI does not certify or endorse any organization’s real-world redaction practices, and a framework saying it will disclose redactions is not the same as verified proof that it does so consistently. The xAI question above is a good illustration of why that caveat matters: NIKOLAI’s own crosswalk flags its own finding as unresolved rather than asserting it as fact, and this guide has kept that framing rather than smoothing it into a firmer claim than the source material supports.
Frequently Asked Questions
Does any framework require a frontier developer to publish the redacted content itself?
No. Every framework discussed here, including SB 53, permits withholding the underlying information — the disclosure duty, where one exists, covers the fact that a redaction happened and a justification for it, not the redacted material itself. California’s SB 53 is the outlier in also requiring five years of unredacted retention, creating a paper trail a regulator or court could eventually reach even though the public never sees it.
Is OpenAI’s approach non-compliant with anything?
Not based on the frameworks compared here. OpenAI’s Preparedness Framework is a voluntary company policy, not a statute, and its text permits redaction without committing to a matching public disclosure duty. That is a difference in what OpenAI has chosen to commit to on paper, not a violation of an external requirement; OpenAI is not shown here as bound by SB 53 or by the EU GPAI Code’s specific redaction-disclosure Measures for purposes of this comparison.
Does the EU’s GPAI Code allow more redaction than the US frameworks?
Not necessarily — read the other way, the EU Code’s regulator-facing default (unredacted access to the AI Office, with only narrow carve-outs) is arguably stricter than what any lab commits to for its own public reports. What the EU Code lacks, on the text reviewed here, is a requirement that a signatory tell the public, specifically, that a redaction occurred and why — the kind of duty Anthropic’s, Meta’s, and California’s text spell out.
Where does this comparison come from?
Every quotation above is drawn directly from the named framework’s own published text — Anthropic’s Responsible Scaling Policy v3.4, OpenAI’s Preparedness Framework v2, Meta’s Advanced AI Scaling Framework v2, California SB 53 as codified, the EU’s GPAI Code of Practice (Safety and Security Chapter), and METR’s own incident writeup — not from secondary summaries. The structure follows NIKOLAI’s Redaction element, CASRAI’s own independent crosswalk.







