Skip to main content
v2026.11,858 entries · CC-BY 4.0

Risk Domain: Why ‘AI R&D’ Lives in Four Different Places

Harmful manipulation is named but unaddressed in Anthropic’s, OpenAI’s, and Meta’s own frameworks — and xAI’s current framework names it as a domain, then never writes the section. Meanwhile “AI R&D” is classified four structurally different ways across four labs. NIKOLAI’s Risk Domain element (N1) maps all nine.

Written and maintained by CASRAI Editorial Board

Last updated

Read the fine print of nine frontier AI safety frameworks and one category keeps showing up unfinished. Anthropic’s Responsible Scaling Policy names four threat models — but “harmful manipulation” isn’t one of the four; it isn’t in the RSP’s own threat-model list at all. OpenAI’s Preparedness Framework v2 has exactly three Tracked Categories, and persuasion is explicitly routed elsewhere: “Persuasion risks will be handled outside the Preparedness Framework.” Meta’s Advanced AI Scaling Framework v2 defines three catastrophic-risk domains — Chemical & Biological, Cybersecurity, Loss of Control — and harmful manipulation is not a fourth. And xAI’s own current Frontier Artificial Intelligence Framework names harmful manipulation as one of “four primary risk domains” in its introduction, then writes a lettered “Addressing” subsection for three of them — (a) CBRN, (b) Loss of Control, (c) Offensive Cybersecurity — and simply never gets to a fourth.

That’s the first thing a side-by-side reading of these documents turns up. The second is stranger: “AI R&D” — a model automating or accelerating its own development — doesn’t have a stable home anywhere in this literature. OpenAI tracks it as its own named category. Google DeepMind folds it into a domain shared with model misalignment. Anthropic treats it as a threat model, not a risk domain, with its own capability threshold. And Meta doesn’t give it a domain at all — it appears as a named enabling capability one level down, filed under Loss of Control. Four organizations, four structurally different answers to where the riskiest capability in the room belongs.

NIKOLAI, CASRAI’s own independent, unendorsed reference dictionary for frontier AI safety, calls the thing these nine documents are all reaching for Risk domain: the top-level catastrophic-harm category that threat models, thresholds, and evaluations get grouped under. It’s element N1 in NIKOLAI’s Actors, Models, and Scope track. This is a primary-source reading of what nine organizations actually wrote, checked against NIKOLAI’s own published crosswalk for the element — not a claim that any of them would describe their own document this way.

What “risk domain” is supposed to organize

NIKOLAI’s own definition frames risk domain as “the top-level category of catastrophic harm under which threat models, thresholds and evaluations are grouped, drawn as a controlled ladder of values (e.g. CBRN, Cyber offence, Loss of control, Harmful manipulation).” In practice, every framework examined here is trying to answer the same prior question before any capability threshold or safety case gets written: which large buckets of harm are we even organizing our safety work around? The answer nine organizations gave is four names, repeated with real but inconsistent overlap — and, in the case of two capabilities in particular, no consistent placement at all.

Anthropic: four threat models, and “automated R&D” is one of them — not a domain

Anthropic’s Responsible Scaling Policy doesn’t use “risk domain” as its own organizing term; its four named threat models cover biological weapons, offensive cyber operations, loss of control, and automated AI research and development, the last tied to a specific capability threshold the RSP calls “AI R&D-4” — defined around “the ability to fully automate entry-level AI research work, and the ability to cause dramatic acceleration in the rate of effective scaling.” That’s a structural choice worth naming plainly: Anthropic’s own document treats automated R&D as a threat model, on the same list as CBRN and cyber, rather than as a domain those threat models sit under. Reviewed directly for this piece, the RSP does not independently list a harmful-manipulation or persuasion threat model alongside the four it does name.

OpenAI: “AI Self-improvement” is a Tracked Category; persuasion is explicitly sent elsewhere

OpenAI’s Preparedness Framework v2 names three “Tracked Categories” — “established areas where we have mature evaluations and ongoing safeguards” — and states them by name: “Biological and Chemical capabilities, Cybersecurity capabilities, and AI Self-improvement capabilities.” AI R&D has a category of its own here, distinct from Anthropic’s threat-model framing. A separate, lower tier of “Research Categories” — areas “that do not yet meet our criteria to be Tracked Categories” — currently lists Long-range Autonomy, Sandbagging, Autonomous Replication and Adaptation, Undermining Safeguards, and Nuclear and Radiological. Harmful manipulation appears in neither list. The framework is direct about why: “Persuasion risks will be handled outside the Preparedness Framework, including via our Model Spec.” NIKOLAI’s own crosswalk for this element additionally cites a separate “Frontier Governance Framework” document tracking cyber offense, CBRN, harmful manipulation, and loss of control as OpenAI’s approach outside the Preparedness Framework proper — a claim this piece did not independently re-verify against that specific document, flagged honestly rather than repeated as confirmed.

Google DeepMind: harmful manipulation gets a threshold; “AI R&D” is merged with misalignment

Google DeepMind’s Frontier Safety Framework v3.1 is the one document here that maps a harmful-manipulation capability threshold directly onto a security tier: CBRN uplift, cyber uplift, and harmful-manipulation capability thresholds each map to Security Level 2 in the framework’s own table. ML R&D acceleration maps to Security Level 3, and DeepMind’s highest capability threshold — full automation of Google’s own AI research teams — is written up not as its own domain but folded into that same “ML R&D and Misalignment” grouping, for which the framework recommends Security Level 4 “for the frontier AI field as a whole,” not for DeepMind’s own commitment alone. Of the four organizations in this comparison, DeepMind is the only one that structurally merges automated-R&D risk with model-misalignment risk into one combined domain component, rather than giving it either a standalone category or a plain threat-model listing.

xAI: four named domains in the introduction, three “Addressing” sections in the body

xAI’s current Frontier Artificial Intelligence Framework, effective 30 June 2026 — not a draft document; that status applies to an earlier, superseded February 2025 xAI safety document, not this one — opens by naming “four primary risk domains where significant or severe outcomes are most likely to emerge”: CBRN Risks, Offensive Cybersecurity Risks, Loss of Control Risks, and Harmful Manipulation Risks. Section 2.1 then works through those domains one at a time under lettered headings: “(a) Addressing CBRN Risks,” “(b) Addressing Risks of Loss of Control,” “(c) Addressing Risks of Malicious Offensive Cybersecurity Usage.” There is no “(d).” Having named harmful manipulation as one of its four foundational domains on page 2, the document’s own risk-identification section never returns to write the corresponding subsection for it — a gap in the current, effective version of xAI’s framework, not a stale artifact of an old draft.

Meta: no harmful-manipulation domain at all; “AI R&D” surfaces as an enabling capability under Loss of Control

Meta’s Advanced AI Scaling Framework v2 defines exactly three catastrophic-risk domains in its own words: “The Framework currently focuses on catastrophic risks in three areas: Chemical & Biological, Cybersecurity, and Loss of Control.” Harmful manipulation isn’t a fourth domain, a research area, or a named exclusion — it simply isn’t addressed in this document’s risk-domain structure. Automated AI R&D doesn’t get a domain of its own either, but it does appear, specifically: under Loss of Control’s enhanced-evaluation criteria, the framework names “Autonomous AI R&D” as an example enabling capability — “ability to autonomously and consistently complete tasks which are representative of the work of a researcher to progress AI development” — alongside a companion capability the framework separately calls “evaluation awareness.” Of the four organizations compared here, Meta is the only one that files automated R&D not as a domain, category, or threat model, but as a named enabling capability one level below Loss of Control.

The taxonomy several others echo: the EU’s four specified systemic risks

The EU AI Act’s General-Purpose AI Code of Practice, Appendix 1.4, specifies exactly four systemic risks by name: Chemical, biological, radiological and nuclear; Loss of control; Cyber offence; and Harmful manipulation, the last defined as “the strategic distortion of human behaviour or beliefs by targeting large populations or high-stakes decision-makers through persuasion, deception, or personalised targeting.” That is the same four-way split — CBRN, cyber, loss of control, harmful manipulation — that NIKOLAI’s own definition text uses as its illustrative ladder, and that xAI’s introduction and DeepMind’s threshold table both echo in their own language. Notably, the GPAI Code does not treat “capabilities to automate AI research and development” as a systemic risk domain at all: it appears one level down, in Appendix 1.3.1’s list of model-capability sources that feed into systemic risk identification — a fifth, structurally distinct placement from all four labs’ own treatments.

California SB 53, and why this row reads Close, not Exact

SB 53’s own statutory language names three harm pathways inside its “catastrophic risk” definition: CBRN assistance to a non-expert, cyberattacks or serious crimes conducted without meaningful human oversight, and loss of a model developer’s control over a model. NIKOLAI’s crosswalk rates this row Close/Medium confidence rather than Exact — correctly, on a direct reading of the statute: SB 53 frames these as pathways to a single defined harm threshold that triggers incident-reporting obligations, not as a standalone, named taxonomy of risk domains the way the EU Code or xAI’s framework do. The substance overlaps closely; the structure doesn’t.

The narrow federal reading, and the industry-consortium taxonomy

Executive Order 14409 and NIST’s associated CAISI guidance name a narrower slice of this same territory: advanced cyber capabilities specifically, without a companion CBRN, loss-of-control, or harmful-manipulation category in the same document. The Frontier Model Forum’s Risk Taxonomy and Thresholds report, by contrast, names three domains close to the labs’ own framings — CBRN Threats, Advanced Cyber Threats, and Advanced Autonomous Behavior Threats — in Section 2.4, on page 9 of the report.

Four organizations, four homes for “AI R&D”

Organization Document Where “AI R&D” sits
OpenAI Preparedness Framework v2 Its own named Tracked Category (“AI Self-improvement”), alongside Bio/Chem and Cybersecurity
Google DeepMind Frontier Safety Framework v3.1 Merged into one domain component with model misalignment (“ML R&D and Misalignment”)
Anthropic Responsible Scaling Policy A threat model with its own capability threshold (AI R&D-4) — explicitly not framed as a domain
Meta Advanced AI Scaling Framework v2 A named enabling capability (“Autonomous AI R&D”) one level under the Loss of Control domain, not a domain of its own

Read together, that’s not four labs disagreeing about how dangerous automated AI R&D is. It’s four labs structuring their own safety documents around four genuinely different ideas of what kind of thing the risk is — a tracked capability area, a merged domain, a threat model, or a sub-capability of a different domain entirely. A reader comparing “how does Lab X handle AI R&D risk” across any two of these documents is comparing answers to differently shaped questions.

Nine organizations, one element: NIKOLAI’s crosswalk

Organization Framework Mapping Confidence
Anthropic Responsible Scaling Policy Four threat models: biological weapons, offensive cyber ops, loss of control, automated R&D Exact / High
OpenAI Preparedness Framework v2 & Frontier Governance Framework Bio/Chem, Cybersecurity, Self-improvement (Tracked Categories); FGF said to track cyber offense, CBRN, harmful manipulation, loss of control Exact / High
Google DeepMind Frontier Safety Framework v3.1 CBRN, Cyber, Harmful Manipulation, ML R&D and Misalignment Exact / High
xAI Frontier Artificial Intelligence Framework (30 Jun 2026) CBRN, Offensive Cybersecurity, Loss of Control, Harmful Manipulation risks Exact / High
Meta Advanced AI Scaling Framework v2 Chemical & Biological, Cybersecurity, Loss of Control Exact / High
EU GPAI Code of Practice, Appendix 1.4 Four specified systemic risks: CBRN, Loss of control, Cyber offence, Harmful manipulation Exact / High
California SB 53 CBRN assistance, cyberattacks/serious crimes without oversight, loss of developer control Close / Medium
US Government Executive Order 14409 / NIST CAISI Advanced cyber capabilities focus only Narrow / Medium
Frontier Model Forum Risk Taxonomy and Thresholds, Section 2.4 (p.9) CBRN Threats, Advanced Cyber Threats, Advanced Autonomous Behavior Threats Exact / High

What NIKOLAI maps, and what it doesn’t

CASRAI’s own NIKOLAI project — an independent, unendorsed reference dictionary for frontier AI safety, not a standard any lab or regulator has agreed to — formalizes this comparison as element N1: Risk domain, part of Track N1, Actors, Models, and Scope. Every row in the table above is what NIKOLAI calls a shadow mapping: CASRAI’s own reading of a published document against NIKOLAI’s element definition, not a confirmation from the mapped organization that this is how it would describe itself. None of the nine has filed a formal Mapping Declaration accepting or correcting how NIKOLAI has read their document.

Worth noting as a genuine piece of CASRAI’s own editorial history, not an external fact: NIKOLAI’s internal placement of this element was itself unsettled for a time. Whether “risk domain” belonged in Track N1, as a scope-defining element, or in Track N2 alongside threat models and risk pathways, was disputed inside CASRAI’s own editorial process before being resolved on 2026-09-19 — risk domain stays in N1, kept distinct from N2’s threat-model elements, on the reasoning that a risk domain is the category a threat model gets grouped under, not a threat model itself. That distinction is exactly what this piece found doing the comparison: Anthropic’s own RSP makes the same move, treating automated R&D as a threat model rather than a domain — NIKOLAI’s N1/N2 boundary and Anthropic’s own document draw the line in the same place.

This element sits in the same N1 track as the scope test each of these organizations also has to clear before risk domain even becomes relevant — covered in CASRAI’s separate Who Counts as ‘Frontier AI’? Eleven Scope Tests Compared. And California SB 53 is worth reading across two NIKOLAI element pages, not one: this piece rates SB 53 Close/Medium confidence against Risk Domain, but NIKOLAI’s separate Incident type element (Track N7, Incidents) rates SB 53 Exact/High confidence against its own crosswalk, citing the statute’s four explicit incident categories at 22757.11(d) directly. That’s not an inconsistency in NIKOLAI’s own mapping — it’s a genuine, citable illustration of the difference these two elements are drawing: SB 53 doesn’t write a risk-domain taxonomy, so it reads as a partial match there, but it does write a precise, enumerated incident taxonomy, so it reads as an exact match against Incident type. No standalone guide exists yet for NIKOLAI’s N7 Incident Type element; its own element page is the primary source for that row.

For the full N1–N10 track structure, see NIKOLAI’s track system; for what NIKOLAI is and isn’t, see What Is NIKOLAI? For companion element deep-dives already live, see Capability Thresholds (N3), Security Level (N6), and Evaluator Independence (N8). For the underlying frameworks themselves, see CASRAI’s guides on frontier AI labs’ safety frameworks, the California SB 53 Frontier AI Transparency Act, the EU AI Act GPAI Code of Practice, and the Frontier Model Forum.

Frequently asked questions

Is “harmful manipulation” really missing from Anthropic’s, OpenAI’s, and Meta’s own safety frameworks?

It’s absent from each document’s own risk-domain or threat-model list, on a direct reading of the current published text. Anthropic’s RSP names four threat models and harmful manipulation isn’t one of them. OpenAI’s Preparedness Framework v2 explicitly routes persuasion risk to its Model Spec instead, outside the framework. Meta’s Advanced AI Scaling Framework v2 names three domains, none of them harmful manipulation. Google DeepMind and xAI both do name it — DeepMind maps a capability threshold to it directly; xAI names it as one of four domains but doesn’t write the corresponding “Addressing” subsection for it in the current version of its framework.

Where does “AI R&D” risk actually live across these frameworks?

Four different places. OpenAI gives it a named Tracked Category of its own. Google DeepMind merges it into one domain with model misalignment. Anthropic treats it as a threat model with its own capability threshold, explicitly distinct from a domain. Meta doesn’t give it a domain at all — it surfaces as a named enabling capability under Meta’s Loss of Control domain. None of the four structures match.

Why does California SB 53 get a different confidence rating here than on NIKOLAI’s Incident Type element?

Because the two elements are testing for different things in the same statute. Risk Domain looks for a standalone taxonomy of catastrophic-harm categories, and SB 53 doesn’t write one — it frames three harm pathways inside its incident-reporting trigger, which reads as a Close/Medium-confidence match. Incident Type looks for a discrete incident taxonomy, and SB 53’s 22757.11(d) enumerates exactly that — four named incident categories — which reads as an Exact/High-confidence match. Both ratings describe the same statute accurately; they’re just answering different questions.

Is xAI’s framework still a draft?

Not the current one. xAI’s Frontier Artificial Intelligence Framework carries an effective date of 30 June 2026 and is not marked as a draft anywhere in its text. An earlier xAI safety document, a February 2025 “Risk Management Framework,” was explicitly labeled a draft; it has since been superseded by the framework this piece examines. The missing harmful-manipulation subsection is a gap in the current, effective document, not a leftover from the older draft.

Is NIKOLAI’s crosswalk an official standard?

No. NIKOLAI is CASRAI’s own, unendorsed reference project. Its crosswalk rows are shadow mappings — CASRAI’s own reading of each organization’s published documents against NIKOLAI’s element definitions — not declarations the mapped organizations have made about themselves.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about Risk Domain: Why ‘AI R&D’ Lives in Four Different Places

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →