Skip to main content
v2026.11,858 entries · CC-BY 4.0

NIKOLAI’s Threat Chain: Actor, Pathway, Model

xAI’s own Frontier AI Framework flags the collision itself: xAI uses “pathway” for a risk domain, while Anthropic, OpenAI, and Meta use the same word for the finer-grained causal route beneath a threat model. NIKOLAI’s N2 track exists to catch exactly this kind of drift — and its three elements read better as one causal chain than as three separate terms.

Written and maintained by CASRAI Editorial Board

Last updated

Last verified: September 20, 2026. xAI’s own Frontier AI Framework does something the other frameworks in this crosswalk don’t: it names the terminology collision itself. Read closely, xAI’s document distinguishes three different things people mean by “pathway” in frontier-AI-safety writing — Anthropic’s eight fine-grained misalignment routes, xAI’s own risk domains, and California SB 53’s three incident-outcome categories — and treats them as genuinely different, not interchangeable. That is a rarer act of terminological self-awareness than it sounds, and it is exactly the kind of drift NIKOLAI’s N2 track was built to catch: xAI uses “pathway” to mean the risk domain itself, while Anthropic, OpenAI, and Meta use the same word for the finer-grained causal route that sits beneath a threat model, one level down.

That single-word divergence is a symptom of a bigger pattern. NIKOLAI, CASRAI’s own independent, unendorsed dictionary of frontier-AI-safety elements, splits this territory into three separate elements — Threat-Actor Profile, Risk Pathway, and Threat Model — inside its N2 track, Threat models and risk framing. Most coverage of frontier-AI risk vocabulary treats those as three thin, standalone definitions. Read together, in order, they aren’t three definitions — they’re one causal chain: who could cause harm (the actor), by what route (the pathway), and what threat model results when a lab decides that route is worth prioritizing over the others it considered. That three-part framing isn’t CASRAI’s invention — it mirrors the N2 track’s own stated scope: “who could misuse the model, through what mechanism, toward what harm.” This page follows that chain link by link, crosswalking each element against the labs, regulators, and standards bodies that publish language in its territory.

What NIKOLAI Means by the Chain

NIKOLAI’s own working definitions — each labeled a proposed, editorial synthesis rather than a quotation from any single source, and each currently tagged proposed in NIKOLAI’s schema (release nikolai-v0.1, inside the broader nikolai-v0.2 dictionary) — read as three links in one chain:

  • Threat-Actor Profile (N2.1): “rates an adversary capability and access on a ladder, from novice individual to nation-state, rather than free text.”
  • Risk Pathway (N2.2): “maps the specific causal route from a model behavior or misuse to harm, detailing the concrete steps a threat model could follow” — finer-grained than a threat model, one level down.
  • Threat Model (N2.3): “states who could cause what harm, by what means, via what role of the AI system, and why that threat was prioritized over others.”

Put plainly: a threat-actor profile answers “who.” A risk pathway answers “by what route.” A threat model is the synthesis — who, by what route, what harm, and why a lab chose to prioritize that combination over every other one it considered. None of the eight sources below use NIKOLAI’s own three-term structure internally; NIKOLAI is reading each source’s published language and classifying it against this chain, which is why every row that follows is what NIKOLAI calls a shadow mapping rather than a confirmed one — CASRAI’s own independent read, not an official standard and not anything the organization named has endorsed.

Link One: Threat-Actor Profile — Five Sources, One Scoping Split

Verified directly against NIKOLAI’s live Threat-Actor Profile element page on September 20, 2026. Five organizations publish language in this element’s territory.

Organization Source Match Confidence What it actually says
Anthropic Risk Report, August 2026 Exact High Defines CB-1 (“individuals/small groups with limited resources”) and CB-2 (“moderately resourced threat actors”), plus “sophisticated insiders” with system access.
OpenAI Preparedness Framework v2 / GPT-5.6 deployment report Close High “Novice actor” with basic technical background through “expert” enabling dangerous novel threats; “spray and pray” by moderately skilled, low-resourced groups.
Google DeepMind Gemini 3.7 Flash Frontier Safety Framework report Exact High TAC-1 (“script kiddie” programmer) through TAC-4 (95th-percentile-plus cybersecurity expert / nation-state groups) on a capability ladder.
xAI Frontier AI Framework, 30 Jun 2026 draft Narrow Medium Identifies “sophisticated non-state actors, insider threats, state-sponsored actors” — a security-focused construct, narrower than the malicious-use profiles used elsewhere.
Meta Advanced AI Scaling Framework v2 Close High Distinguishes state/non-state and high/low-skill actors; covers “low and moderate skill” proliferation scenarios.

There’s a sixth data point that doesn’t fit neatly into the five-row table: the EU’s GPAI Code of Practice, in its Safety and Security chapter, adopts the same narrow, security-only framing xAI does. That leaves a real split down the middle of this element. Anthropic, OpenAI, DeepMind, and Meta each build a threat-actor construct broad enough to cover malicious use of a model generally — a low-skill actor doing something harmful with an ordinary capability counts. xAI and the EU GPAI framework instead scope their threat-actor profiles to security specifically: who could compromise, steal, or subvert the system itself, not who could misuse its ordinary outputs. Both are legitimate ways to define “threat actor” — but a reader who assumes the term means the same thing in xAI’s framework as it does in Anthropic’s is assuming away a real scoping difference.

Link Two: Risk Pathway — Seven Sources, and the False-Friend That Started This Page

Verified directly against NIKOLAI’s live Risk Pathway element page on September 20, 2026. Seven sources publish language in this element’s territory, two more than Threat-Actor Profile — California SB 53 and the Frontier Model Forum both have something to say about pathways that they don’t say about threat actors specifically.

Organization Source Match Confidence What it actually says
Anthropic Risk Report, August 2026 (§2.2.1, §2.13) Equivalent High Eight priority pathways named, including “diffuse sandbagging on safety R&D,” “self-exfiltration,” and “persistent rogue internal deployment.” Its own hedge: “We aren’t able to defend the choice of these pathways rigorously.”
OpenAI Preparedness Framework v2 / Path to Astra Close High Safeguards Report identifies “ways a risk of severe harm can be realized”; Astra defines cyber pathways as either a malicious actor using the model to develop exploits, or the model itself taking unauthorized action.
Google DeepMind Frontier Safety Framework v3.1 Close High Critical capability levels set by “identifying and analyzing the main foreseeable paths through which a model could cause severe harm.”
xAI Frontier AI Framework, 30 Jun 2026 draft No mapping (false friend) High Uses “pathway” to mean risk domains themselves (§2.1) — not a finer-grained causal route. The document’s own false-friends register names the collision directly, distinguishing Anthropic’s eight fine-grained routes, xAI’s risk domains, and SB 53’s three incident categories as three different senses of the word.
Meta Advanced AI Scaling Framework v2 Close High Threat modeling “identifies the potential causal pathways for realizing the catastrophic outcome.”
California SB 53 §22757.11(c) Broader Medium Defines catastrophic risk by three outcome categories (CBRN assistance, autonomous cyberattack or serious crime, loss of control) rather than by naming specific causal pathways — a broader, outcome-defined construct, confirmed against the bill’s own text.
Frontier Model Forum Risk Taxonomy and Thresholds Close High Threat modeling includes “mapping the potential pathways to those outcomes”; sets a credibility bar of “there is a credible pathway to extreme harm.”

CASRAI independently pulled California SB 53’s statutory text directly from the legislature’s own bill tracker to check this row, rather than relying on NIKOLAI’s summary alone. Section 22757.11(c) defines catastrophic risk as a foreseeable, material risk of death or serious injury to more than 50 people, or more than a billion dollars in property loss, from a single incident where a frontier model does one of three things: provides expert-level CBRN weapon assistance, commits an autonomous cyberattack or serious crime with no meaningful human oversight, or evades its developer’s or user’s control. That confirms NIKOLAI’s read: SB 53 legislates around outcomes, not around the specific causal steps that get a model there — which is a genuinely different shape of definition than Anthropic’s eight named pathways or DeepMind’s “foreseeable paths.”

The xAI row is the one that gives this page its name. Most false-friend terminology traps in this cluster’s coverage so far have been ones CASRAI had to notice by comparing documents against each other. This one is different: xAI’s own Frontier AI Framework names the collision itself, distinguishing its own risk-domain sense of “pathway” from Anthropic’s finer-grained causal-route sense and from SB 53’s outcome-category sense, in its own text. That is a genuine, source-confirmed divergence, not an inference CASRAI is drawing across documents that never engaged with each other — and it is exactly why NIKOLAI splits “risk pathway” out as its own element (N2.2) instead of treating it as a synonym for “threat model” or “risk domain.” A reader who sees “pathway” in an xAI document and assumes it means the same granular, step-by-step route Anthropic’s Risk Report maps will misread the framework’s structure.

Link Three: Threat Model — Eight Sources, and Where the Chain Closes

Verified directly against NIKOLAI’s live Threat Model element page on September 20, 2026. Eight sources publish language here — one more than Risk Pathway, because the EU’s GPAI Code of Practice and the Safety Framework Cards working paper both have something to say specifically about how a completed threat model gets structured and disclosed, beyond what either says about actors or pathways alone.

Organization Their term & source Match Confidence What it describes
Anthropic “CB-1 threat model” (Risk Report, Aug 2026, §4.1-6.1) Exact High Individuals/small groups using AI for non-novel CBRN weapon access; prioritized by expected damages, a clear AI-role distinction, historical sanity checks, and generalizability.
OpenAI “Threat model” (Preparedness Framework v2, §2.2) Close High Identifying specific severe-harm risks from frontier capabilities per tracked category, reviewed and approved by its Safety Advisory Group.
Google DeepMind Report method terms (Gemini 3.7 Flash FSF report, p.8) Close High Uses “threat actor type,” “scenario,” “harm journey,” and “bottleneck sub-stages” as its own working vocabulary for the same territory.
xAI “Systemic risk scenarios” (Frontier AI Framework, 30 Jun 2026 draft) None (undeclared) Medium Enumerates causal factors, potential harms, and mitigations, but the document’s own metadata still marks it a confidential working draft.
Meta “Threat modeling” & “threat scenarios” (Advanced AI Scaling Framework v2, Appendix I; §3.4) Exact High A structured process naming how frontier AI could contribute to a specific outcome; scenarios specify threat actors, enabling capabilities, and deployment context, tagged with identifiers like “Cyber 1” and “TS.1.1.”
EU Measure 2.2, “systemic risk scenarios” (GPAI Code of Practice, Safety and Security chapter) Close High Signatories develop systemic risk scenarios per identified risk, feeding a broader systemic-risk model, with no public identifier scheme or mandatory publication beyond the Model Report.
Frontier Model Forum “Threat modeling” & “threat scenarios” (Risk Taxonomy and Thresholds, §2.1, pp.7-8) Exact High A process for systematically anticipating how a threat actor would leverage frontier AI; scenarios specify tasks, exploitable capabilities, and complementary tools.
Safety Framework Cards “Risk ontology” dimension (SSRN working paper 7061798) None (unverified) Low Recorded from a discovery-stage snippet only — the full paper sits behind SSRN’s paywall and was not read in full for this element. Flagged honestly rather than dropped.

CASRAI independently confirmed the EU’s Code of Practice Safety and Security chapter is real, live, and scoped to systemic risk from the most advanced models — checked directly against the European Commission’s own digital-strategy page on September 20, 2026 — though the chapter’s own Measure 2.2 text sits behind a PDF the Commission distributes on request rather than serving inline, so its exact wording is NIKOLAI’s own read rather than something this page re-quotes independently. What both NIKOLAI’s crosswalk and the Commission’s own page agree on is the shape of the requirement: systemic risk scenarios feed a broader risk model, with no public identifier scheme and no disclosure obligation beyond the Model Report a signatory already has to file.

Meta’s Gap: A Named Scenario Isn’t a Published Playbook

The most useful single row in this element belongs to Meta, and it’s useful because of what it deliberately leaves out. Meta’s Advanced AI Scaling Framework v2 assigns real, trackable identifiers to its threat scenarios — “Cyber 1,” “TS.1.1” — the same instinct toward structured, referenceable IDs that shows up across this element’s better-documented rows. But Section 3.4 is explicit that Meta “withholds full details of the constituent steps and tasks within a threat scenario.” The identifier is public. The operational path it labels is not. That is a real, and reasonable, public-summary-versus-restricted-detail split — a lab can be transparent about the fact that it modeled a specific cyber threat scenario without publishing the steps an attacker would need to reproduce it. It also means a reader comparing frameworks side by side sees a named “Cyber 1” entry in Meta’s disclosures and has no way to check, from the public document alone, whether Meta’s Cyber 1 lines up with DeepMind’s cyber uplift threat model or OpenAI’s cyber pathway definition in Astra. The ID system is a genuine step toward comparability; it stops short of making the comparison possible from the outside.

Reading the Chain End to End

Lined up, the three elements answer three different questions that only make sense in sequence. Threat-Actor Profile answers “who” — and the honest finding there is a scoping split, not a vocabulary gap: xAI and the EU GPAI framework build a narrower, security-only “who” than Anthropic, OpenAI, DeepMind, and Meta do. Risk Pathway answers “by what route” — and the honest finding there is xAI’s own admitted collision, where the word that means a fine-grained causal route everywhere else in this crosswalk means a risk domain in xAI’s own framework, by xAI’s own account. Threat Model is the synthesis of both, plus the “why prioritized” question none of the other two elements ask alone — and the honest finding there is Meta’s transparent-ID, opaque-detail split, which is a genuine design choice a lab makes about how much of its own threat model to publish, not a vocabulary failure.

None of that is a criticism of any single organization. A narrower threat-actor scope, an idiosyncratic use of “pathway,” and a withheld operational detail are each defensible choices a safety team can make for good reasons. The problem NIKOLAI’s N2 track exists to solve is narrower than “is everyone doing this right” — it’s that a reader moving between Anthropic’s Risk Report, xAI’s Frontier AI Framework, and Meta’s Advanced AI Scaling Framework needs to know that “threat actor,” “pathway,” and “threat model” don’t automatically mean the same thing in each one, and right now, checking requires reading each source’s own language directly, the way this page just did, one link of the chain at a time.

Where NIKOLAI Fits In

This page is CASRAI’s own deep-dive on all three elements of NIKOLAI’s N2 track, Threat models and risk framing: Threat-Actor Profile (N2.1), Risk Pathway (N2.2), and Threat Model (N2.3). NIKOLAI is not affiliated with, run by, or endorsed by Anthropic, OpenAI, Google DeepMind, xAI, Meta, the European Commission, the State of California, the Frontier Model Forum, or the authors of the Safety Framework Cards working paper, and none of them has been consulted on how NIKOLAI classifies their language. Every row in all three tables above — eighteen rows across three elements — is what NIKOLAI calls a shadow mapping: CASRAI’s own independent reading of a published document, carrying its own confidence label (Exact/Close/Narrow/Broader/None) and evidence tier (High/Medium/Low), and nothing more than that unless the organization in question files an explicit Mapping Declaration confirming how it actually uses the term. As of this writing, none of the eight organizations and bodies named across these three elements has done so.

This page joins two other N2-adjacent element deep-dives already published in this cluster. Risk Domain (N1, in the Actors, Models and Scope track) is the classification layer a threat model’s harm category draws from — CBRN, cyber, loss-of-control, manipulation — and is the exact term xAI’s “pathway” collides with, since xAI uses “pathway” for what other frameworks would call a risk domain. Evaluation-Validity Threats (N5, in the Evidence and Evaluations track) is what a lab runs to check whether a threat model’s assumed capability actually holds up under testing, once the chain this page describes has named who, by what route, and what result. Read together with the two guides that established this element-deep-dive format for the cluster — Capability Thresholds: 14 Labs and Regulators, One Undefined Term and Security Level: Nine Labs and Regulators, One Undefined Standard — the pages trace a fuller thread: a threat model names what a lab is worried about, a capability threshold marks when a model has crossed into that territory, and a security level is supposed to hold once it has.

Frequently Asked Questions

What is NIKOLAI’s N2 track?

N2, “Threat models and risk framing,” is one of NIKOLAI’s ten tracks. It holds three elements — Threat-Actor Profile, Risk Pathway, and Threat Model — that together answer who could misuse a model, through what mechanism, and toward what harm. NIKOLAI is CASRAI’s own independent, unendorsed dictionary of frontier-AI-safety elements; it does not claim any lab adopted its specific three-term structure.

What is the false-friend problem with “pathway”?

xAI’s Frontier AI Framework uses “pathway” to mean a risk domain — the general category of harm, like cyber or CBRN. Anthropic, OpenAI, and Meta use the same word for something narrower: the specific, step-by-step causal route from a model behavior to a harmful outcome, one level more granular than a threat model. xAI’s own document names this collision directly, distinguishing its risk-domain sense from Anthropic’s finer-grained sense and from California SB 53’s outcome-category sense.

Why do xAI and the EU GPAI framework scope “threat actor” more narrowly than other labs?

Both build a security-focused threat-actor construct — who could compromise, steal from, or subvert the AI system itself — rather than the broader malicious-use construct Anthropic, OpenAI, Google DeepMind, and Meta use, which also covers a low-skill actor misusing the model’s ordinary outputs. Both are legitimate scoping choices; they are not the same scope.

What does it mean that Meta “withholds” threat-scenario detail?

Meta’s Advanced AI Scaling Framework v2 assigns public identifiers to its threat scenarios, such as “Cyber 1” and “TS.1.1,” but explicitly states it withholds the constituent steps and tasks within each scenario. The identifier is public; the operational detail behind it is restricted. That is a deliberate public-summary-versus-restricted-detail split, not an omission.

Is NIKOLAI’s crosswalk an official or endorsed mapping?

No. Every row across all three N2 elements is a shadow mapping — CASRAI’s own independent reading of what each organization has published — unless that organization has filed an explicit Mapping Declaration confirming it. As of this writing, none of the organizations and bodies named in this crosswalk has done so.

Related Reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about NIKOLAI’s Threat Chain: Actor, Pathway, Model

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →