Skip to main content
v2026.11,858 entries · CC-BY 4.0

AI Safety vs. AI Security: What the Distinction Actually Means

A short definitional guide distinguishing AI safety (preventing an AI system from causing unintended harm through its own behavior or capabilities) from AI security (protecting AI systems from external attack or misuse, such as model weight theft, adversarial attacks, and data poisoning), grounded in how frontier labs and oversight institutions actually use the terms.

Written and maintained by CASRAI Editorial Board

Last updated

“AI safety” and “AI security” get used interchangeably in casual conversation, but inside the frontier labs and the institutions that oversee them, they name two different disciplines with different failure modes, different teams, and different mitigations. Confusing the two leads to gaps: a lab can pass every security audit and still ship a model that causes harm through its own behavior, and a model can be behaviorally well-aligned while its weights sit exposed to theft.

The short version

AI safety is concerned with an AI system causing unintended harm through its own behavior or capabilities — the model does something its developers did not want, whether through misalignment, a dangerous capability, or an inadequate deployment safeguard. AI security is concerned with protecting an AI system from external attack or misuse — someone else stealing the model, manipulating its training data, or exploiting it to extract information or bypass controls. Safety is about what the system does; security is about what is done to the system, or what someone else does with it once compromised.

What AI safety covers

This is the focus of most of the frameworks covered elsewhere in this cluster. Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework, and Google DeepMind’s Frontier Safety Framework are all, at their core, safety documents: they define capability thresholds, run dangerous-capability evaluations (CBRN uplift, cyberattack automation, AI R&D acceleration), and assign safeguard tiers that scale with what a model can do. See Responsible Scaling Policy (RSP): what it is and how the major labs compare for how those thresholds and tiers work in practice.

Safety work asks questions like: does this model pursue the goals its developers intended, including in situations they didn’t anticipate? Will it refuse harmful requests reliably? Could its capabilities, if deployed without safeguards, materially assist someone attempting mass-casualty harm? These are questions about the model’s own behavior and capability profile, not about an outside attacker.

What AI security covers

Security work protects the model and its supporting infrastructure from people trying to compromise it. Anthropic’s own account of its security program frames this directly: frontier model security means protecting “advanced models and model weights, and the research that feeds into them” from “theft or misuse,” and the company argues frontier AI research “must be secured to levels far exceeding standard practices for other commercial technologies.” That is a security claim, not a safety claim — it is about controlling who can access a set of trained weights, not about what the model does when it runs.

NIST’s taxonomy of adversarial machine learning (NIST AI 100-2) groups the standard attack categories this work defends against: evasion attacks, where an attacker crafts inputs designed to fool a deployed model at inference time; poisoning attacks, where an attacker corrupts training or fine-tuning data so the resulting model behaves incorrectly or contains a hidden backdoor; and extraction attacks, where an attacker queries a model to reconstruct its parameters, training data, or functionality without authorization. Model weight theft — an outside party exfiltrating the trained weights themselves — sits alongside these as the highest-value target, since possessing the weights bypasses every other control at once.

The institutional naming makes the split explicit: the body the UK government stood up to evaluate frontier models is called the AI Security Institute, not the AI Safety Institute — a deliberate rename reflecting that its remit spans both misuse and loss-of-control risks. See CAISI and the UK AI Security Institute: how pre-deployment testing agreements work for how that institute’s work relates to the US Center for AI Standards and Innovation (CAISI).

Where the line gets blurry

The two disciplines aren’t hermetically sealed. A poisoned training set is a security failure that produces a safety-relevant outcome (a model that behaves unsafely). A model with weak refusal behavior can be jailbroken by a determined user — is that a safety gap (the model should refuse regardless) or a security gap (the attacker circumvented a control)? In practice, labs treat it as both: safety teams own the underlying behavior, security teams own who can reach the model and with what access.

NIST’s AI Risk Management Framework is one place this shows up structurally: it lists “safe” and “secure and resilient” as separate characteristics of a trustworthy AI system, rather than folding one into the other — an explicit acknowledgment that a system can satisfy one without satisfying the other.

Safety, security, and alignment: three different questions

A third term compounds the confusion. Alignment is not a synonym for safety or security — it is a technical sub-problem inside safety. Alignment asks whether a model’s own goals and behavior track what its developers intended; safety is the umbrella covering alignment alongside other sub-problems such as adversarial robustness, interpretability, and capability evaluation. Security is a separate discipline again, concerned with external threats rather than the model’s internal objectives. See What is AI alignment? Definition and why it matters for frontier AI safety for the full breakdown of how alignment fits under the safety umbrella, and how it differs from both adjacent terms.

A useful shorthand: a model can be well-aligned and still insecure (its weights sit unprotected), and a model can be extremely secure and still unsafe (nothing external threatens it, but its own behavior is unreliable). The three properties are independent enough that a framework addressing only one of them is not addressing the others.

Why this distinction is worth tracking precisely

Because these terms map to different named frameworks, different regulatory hooks, and different reporting obligations. California’s SB 53, for example, imposes critical safety incident reporting requirements that are triggered by safety-relevant events — not by every security incident. A lab that conflates the two in its public communications, or a research administrator reading a lab’s framework and treating “security” claims as evidence of “safety” work (or vice versa), risks drawing the wrong conclusion about what a framework actually covers. CASRAI’s NIKOLAI dictionary exists partly to prevent that kind of drift: it assigns stable, versioned definitions to terms like these across the frameworks published by different labs, so a claim made under one lab’s policy can be compared against another’s using the same vocabulary rather than each lab’s own framing. See NIKOLAI, CASRAI’s dictionary of elements for frontier AI safety documentation.

FAQ

Is model weight theft a safety issue or a security issue?

Security. Weight theft is an external attacker gaining unauthorized access to trained model parameters. It becomes safety-relevant downstream — a stolen model with safety guardrails removed could be misused — but the theft itself is a security failure, and labs’ own frameworks (Anthropic’s security program, for instance) treat it as such.

Does NIST treat AI safety and AI security as the same thing?

No. NIST’s AI Risk Management Framework lists “safe” and “secure and resilient” as distinct characteristics of a trustworthy AI system, and its separate adversarial machine learning taxonomy (AI 100-2) addresses security-specific attack categories — evasion, poisoning, and extraction — independently of safety concepts like capability evaluation.

Where does alignment fit relative to safety and security?

Alignment is a sub-problem inside safety, not a separate third category alongside it. It specifically concerns whether a model’s behavior matches its developers’ intent. Security is the separate discipline of protecting the model from external interference. See What is AI alignment? for detail.

Why do some institutions use “safety” in their name and others use “security”?

The naming is not fully standardized across governments and labs, and it has shifted over time — the UK’s institute was renamed from “AI Safety Institute” to “AI Security Institute” to better reflect a remit that spans both misuse and loss-of-control risks alongside safety evaluation. Treat institutional names as a starting point, not a precise scope statement, and check what each body’s mandate actually covers.

Can a framework address only one of the two?

Yes, and many published frameworks lean heavily toward one side. Capability-threshold documents like the RSP, Preparedness Framework, and Frontier Safety Framework are primarily safety documents; a lab’s security whitepaper covering weight protection and infrastructure controls is a separate document with a separate scope. Reading only one gives an incomplete picture of a lab’s overall risk posture.

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about AI Safety vs. AI Security: What the Distinction Actually Means

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →