Skip to main content
v2026.11,858 entries · CC-BY 4.0

The One Redaction Reason No Lab Will Admit To

Anthropic, SB 53, Meta, OpenAI, the EU’s GPAI Code, and METR each name reasons a frontier AI developer can redact for. Together they cover six categories — but none of them, in its own words, rules out the one reason everyone would call illegitimate. Built from NIKOLAI’s Redaction reason element (N8.6), distinct from the Redaction element covered in CASRAI’s companion disclosure guide.

Written and maintained by CASRAI Editorial Board

Last updated

Six ways to justify a cut, and zero ways to admit the real one

Every frontier-AI safety framework that discusses redaction lists reasons a developer can cite for cutting material from a published report. Line the lists up — Anthropic’s Responsible Scaling Policy, California’s SB 53, Meta’s Advanced AI Scaling Framework, OpenAI’s Preparedness Framework, the EU’s GPAI Code of Practice, and METR’s own incident writeups — and a six-category vocabulary emerges: legal compliance, intellectual-property protection, public safety, privacy, cybersecurity, and national security. No single source names all six; each contributes a subset, and only by reading them together does the full list appear.

What none of the six sources does, on the text available, is write down the reason everyone would presumably agree is illegitimate: cutting a finding because it makes the developer look bad. That gap — a taxonomy built entirely out of permitted categories, with the prohibited case left unstated in every framework’s own words — is what this guide is about. It is built from NIKOLAI’s Redaction reason element (N8.6), a narrower and more recent addition to Track N8 than the Redaction element (N8.7) CASRAI’s companion guide on what gets redacted and who has to say so already covers.

Redaction reason is not the same element as Redaction

NIKOLAI keeps these as two separate, adjacent elements inside Track N8 (Transparency and Review), and the distinction matters for what each one is checking. The Redaction element (N8.7) is about the event itself — CASRAI’s own editorial proposal that a real redaction record should capture where in the document something was cut, roughly how much, who decided, and that the cut is disclosed rather than silent. CASRAI’s companion guide builds its whole comparison around that element: which frameworks commit to telling anyone a redaction happened at all.

The Redaction reason element (N8.6) assumes a redaction has already been disclosed and asks a narrower question: what value goes in the “why” field, and where is the line between a legitimate reason and an illegitimate one. NIKOLAI’s own definition for N8.6 states that a controlled vocabulary for this field “should include permitted reasons (legal compliance, intellectual-property protection, public-safety considerations, privacy, cybersecurity, national security) and prohibited reasons — grounds developers cannot cite, such as redacting unfavorable findings.” That sentence is CASRAI’s own vocabulary design, written to describe what a complete field should look like — it is not a quotation from Anthropic, SB 53, or any other source, and this guide is not going to present it as one. What each framework actually says, on its own terms, is below.

The permitted reasons, framework by framework

Anthropic’s Responsible Scaling Policy, version 3.4, states its list directly, in section 3.5 (Publication and Redactions): “Reasons we may redact material include but are not limited to” four named categories — “Legal compliance,” “Intellectual property protection,” “Public safety considerations,” and “Privacy.” The “include but are not limited to” framing is worth noting on its own: even Anthropic’s list is explicitly open, not a closed set of exactly four reasons.

California’s SB 53 — the Transparency in Frontier Artificial Intelligence Act — comes closest to naming all six in one place. Section 22757.12(f)(1) permits redaction “necessary to protect the frontier developer’s trade secrets, the frontier developer’s cybersecurity, public safety, or the national security of the United States or to comply with any federal or state law.” That single sentence supplies trade secrets (intellectual property), cybersecurity, public safety, national security, and legal compliance — five of the six categories NIKOLAI’s vocabulary names, missing only privacy. CASRAI’s separate guide on SB 53’s accountable-decision-maker requirements covers the statute’s governance side; the redaction-and-retention clause at 22757.12(f)(2) is covered in the disclosure guide linked above.

Meta’s Advanced AI Scaling Framework, version 2, states a narrower pair in section 2.2.1: it will redact “to protect trade secrets or as appropriate under law” — intellectual property and a general legal-compliance catch-all, with no separate mention of public safety, privacy, cybersecurity, or national security as named categories in that sentence. OpenAI’s Preparedness Framework, version 2, permits redaction “such as to protect intellectual property or safety” — two categories, stated as examples rather than an exhaustive list, and without OpenAI committing to a matching disclosure duty the way Anthropic’s and Meta’s text do.

The EU’s GPAI Code of Practice takes a different shape again: rather than listing reasons a signatory may redact for the public, its Measures name carve-outs for what can be withheld from the regulator. Measure 7.7 permits withholding Model Report material from the AI Office only when “required by national security laws,” and Measure 9.2 permits redacting serious-incident information “to comply with other Union law.” National security and legal compliance are the only two categories the Code’s regulator-facing text names; it does not separately name intellectual property, public safety, privacy, or cybersecurity as redaction grounds in the Measures reviewed for this guide.

METR sits outside this list rather than inside it, because METR is not stating its own reasons for redacting — it is vouching for a report it evaluated but did not write. Its incident writeup on an OpenAI/Hugging Face investigation opens with: “Except where explicitly noted in this post, OpenAI redacted no additional information that was important to our conclusions.” That is a completeness attestation, not a reason taxonomy, and this guide is not treating it as a sixth first-party source for the permitted-reasons list.

The reason nobody writes down

Read across all six sources, no framework’s own text states a prohibited reason — there is no sentence in Anthropic’s RSP, SB 53, Meta’s AASF, OpenAI’s Preparedness Framework, or the EU’s GPAI Code that says, in effect, “a developer may not redact a finding merely because it is unfavorable.” The closest any of the source text reviewed for this guide comes is a footnote in Anthropic’s RSP, in the section describing what a redaction-focused external reviewer checks: “we may redact information if sharing it would endanger our legal right to treat it as confidential. This would be an uncommon situation. It does not mean we would redact information merely because it is sensitive or to reduce the risk of a leak.” That rules out two specific bad reasons — sensitivity alone, and leak-avoidance — but it is not the same sentence as ruling out redacting an unfavorable finding, and this guide is not treating it as one.

What Anthropic’s RSP does build, instead of a written prohibition, is a structural check aimed at the same outcome. Section 3.6.3 has the external reviewer specifically evaluate “redaction justification” (whether the reviewer agrees with the publicly stated reasons) and “materiality” (whether a redaction is material to any of the reviewer’s own disagreements with the report’s conclusions) — a mechanism built to surface exactly the case where a redaction is covering up an inconvenient result, without the policy needing to name that case in so many words. SB 53’s five-year unredacted-retention duty (22757.12(f)(2)) works the same way from a different angle: it does not forbid an outcome-motivated redaction directly, but it guarantees a regulator or court can eventually reach the unredacted version if one is suspected.

So the honest version of this taxonomy has a gap in it, and NIKOLAI’s N8.6 element is the one place, across everything reviewed for this guide, that names the gap explicitly — as CASRAI’s own proposed field design, flagging what a complete redaction-reason vocabulary would need to rule out, not as a promise any of the six organizations above has made. That is exactly what a shadow mapping is for: NIKOLAI does not need an organization’s permission to notice that its published reasons, added together, still leave the obvious bad reason unaddressed.

How this differs from CASRAI’s redaction-disclosure guide

CASRAI already has a guide answering a related but different question: which frameworks require a developer to disclose that a redaction happened at all, built from the N8.7 Redaction element’s nine-organization crosswalk. That guide’s finding is about disclosure duty — Anthropic, Meta, SB 53, and METR commit to telling someone a redaction occurred and roughly why; the EU’s GPAI Code and what can currently be verified of the UK’s approach do not carry an equivalent public-facing duty.

This guide assumes disclosure already happened and asks a narrower, vocabulary-level question: once a developer says “we redacted this, and here’s why,” what values is “why” actually allowed to take, and which value is everyone implicitly agreeing not to use. The two guides share source material — the same RSP section, the same SB 53 clause — but read it for different things, which is why NIKOLAI keeps them as separate elements rather than folding N8.6 into N8.7. Read the disclosure guide for whether a redaction gets flagged; read this one for what the flag is allowed to say.

NIKOLAI’s Redaction reason element

This guide is built from NIKOLAI’s Redaction reason element (N8.6), part of Track N8 — Transparency and Review. NIKOLAI is CASRAI’s own frontier-AI-safety dictionary — an independent, unendorsed reference work, not a standard any of the organizations discussed above has adopted. Every mapping in N8.6’s crosswalk is what NIKOLAI calls a shadow mapping: CASRAI’s own reading of a document each organization actually published, not a reading any of them has reviewed or confirmed. None of the six has filed a Mapping Declaration stating how it uses the term “redaction reason” internally, and a framework naming four or five permitted categories on paper is not verified proof of what any specific redaction in any specific report was actually for.

This guide is part of CASRAI’s Regulation & Standards subcluster within the Frontier AI Safety & Governance content cluster. For the full 64-element vocabulary this guide draws two elements from, see the NIKOLAI dictionary and its track-by-track map, N1 through N10.

Frequently asked questions

Is this the same comparison as CASRAI’s other redaction guide?

No. CASRAI’s redaction-disclosure guide asks whether a framework requires telling anyone a redaction happened. This guide assumes that question is settled and asks what reasons the “why” field is allowed to hold once disclosure occurs. NIKOLAI models them as two elements, N8.7 and N8.6, for exactly that reason.

Does any framework explicitly ban redacting an unfavorable finding?

Not in the source text reviewed for this guide. Anthropic’s RSP, SB 53, Meta’s AASF, OpenAI’s Preparedness Framework, and the EU’s GPAI Code each name permitted reasons; none of them states a prohibited one. The closest textual proxy, a footnote in Anthropic’s RSP ruling out redacting “merely because it is sensitive” or “to reduce the risk of a leak,” is narrower than a ban on outcome-motivated redaction. NIKOLAI’s N8.6 element names the prohibited case as CASRAI’s own proposed field value, not as a quotation from any of the six sources.

Which framework names the most permitted-reason categories?

California’s SB 53, in one sentence at 22757.12(f)(1): trade secrets, cybersecurity, public safety, national security, and compliance with federal or state law — five of NIKOLAI’s six categories. Anthropic’s RSP names four (legal compliance, intellectual property, public safety, privacy), explicitly as a non-exhaustive list. No single source names all six; privacy appears only in Anthropic’s text among the sources reviewed here.

Where does this comparison come from?

Every quotation above is drawn from the named framework’s own published text — Anthropic’s Responsible Scaling Policy version 3.4, California SB 53 as codified at Government Code section 22757.12, Meta’s Advanced AI Scaling Framework version 2, OpenAI’s Preparedness Framework version 2, the EU’s GPAI Code of Practice, and METR’s own incident writeup — not from secondary summaries. The structure follows NIKOLAI’s Redaction reason element, CASRAI’s own independent vocabulary proposal.

Related reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about The One Redaction Reason No Lab Will Admit To

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

AI policy question? Get an answer citing the framework.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.