Skip to main content
v2026.11,858 entries · CC-BY 4.0
Frontier AI Safety & Governance

Safety Framework Fundamentals

The foundational vocabulary and cornerstone explainers for frontier-AI safety frameworks: Responsible Scaling Policies, Preparedness Frameworks, Frontier Safety Frameworks, and the concepts (capability thresholds, safety cases, capability elicitation, dangerous-capability evaluations) they define. The catch-all entry point for this cluster.

Guides

Chain-of-Thought Monitorability as a Safety Control

Reasoning models think in readable English, and reading that trace catches reward hacking, scheming and alignment faking that the final output hides. But monitorability is a property no one has guaranteed will persist: optimisation pressure teaches obfuscation, faithfulness is already low, GPT-6 Astra’s own system card reports decreased monitorability, and one lab lost it by accident in ~8% of RL episodes. What the evidence shows, what the rules require (almost nothing), and what an institution can actually monitor through a commercial API.

Emergent Abilities: Are AI Capability Jumps Forecastable?

Frontier safety frameworks assume you can see a dangerous capability coming. The emergent-abilities literature tests that assumption: sharp benchmark jumps are often a metric artifact, but predictability is a separate question and downstream forecasting remains hard. What five forecasting methods have actually demonstrated, how safety frameworks compensate with evaluation cadence, and why the EU AI Act reaches for a FLOP number instead.

Likelihood Term: Three Frameworks, No Shared Probability Scale

Anthropic’s RSP, Google DeepMind’s FSF, and the EU’s GPAI Code all use likelihood language — ‘plausible,’ ‘unlikely’ — without ever defining a probability range, unlike the IPCC’s calibrated scale. NIKOLAI’s Likelihood Term (N4) proposes the fix.

Marginal Risk vs. Absolute Risk: The Contested Framing Behind Frontier AI Safety Cases

Anthropic’s RSP rewrote itself around a distinction between a model’s absolute risk and its marginal risk relative to what other developers already field — and built in extra governance friction specifically for when that comparison drives a decision. OpenAI, Google DeepMind, and Meta all reason the same way, but each anchors the comparison to a different baseline. This guide sets out what each framework actually says, in its own words, and where the comparison holds up under scrutiny.

Silent Revision: A New Study Finds Most Safety-Framework Changes Go Undisclosed

A frozen codebook, 710 traced commitment instances, and twelve developers’ own version history: a new study puts a number on how much of a safety framework’s real change history never makes it into the developer’s own account of what changed.

The IDAIS-Beijing Statement, Explained: AI Safety’s Red Lines From China

The IDAIS-Beijing Consensus Statement on Red Lines, co-signed by Andrew Yao, Zeng Yi and Xue Lan alongside Bengio, Hinton and Russell, is CASRAI’s first frontier-ai-safety guide to center a non-Western academic-institution voice — and its self-modification and R&D-budget asks map to two specific NIKOLAI elements.

Is AI Safety a Property of the Model? The Narayanan-Kapoor Critique

Arvind Narayanan and Sayash Kapoor’s March 2024 essay ‘AI Safety Is Not a Model Property’ argues that capability-threshold frameworks — Anthropic’s RSP, OpenAI’s Preparedness Framework, Google DeepMind’s FSF — rest on a category error: safety depends on deployment context, not a property a model can be tested for and certified to have. This guide walks through their argument, in full attribution, as the academic counterpoint this cluster’s framework-explainer and comparison pages have been missing.

NIKOLAI’s Threat Chain: Actor, Pathway, Model

xAI’s own Frontier AI Framework flags the collision itself: xAI uses “pathway” for a risk domain, while Anthropic, OpenAI, and Meta use the same word for the finer-grained causal route beneath a threat model. NIKOLAI’s N2 track exists to catch exactly this kind of drift — and its three elements read better as one causal chain than as three separate terms.

Risk Domain: Why ‘AI R&D’ Lives in Four Different Places

Harmful manipulation is named but unaddressed in Anthropic’s, OpenAI’s, and Meta’s own frameworks — and xAI’s current framework names it as a domain, then never writes the section. Meanwhile “AI R&D” is classified four structurally different ways across four labs. NIKOLAI’s Risk Domain element (N1) maps all nine.

Commitment: Ten Labs and Regulators, One Inconsistent Promise

Anthropic and Google DeepMind both publish standards they say, in their own text, they cannot unilaterally commit to meeting. That is what NIKOLAI’s Commitment element is built to catch. Ten labs, regulators, and signatory groups use the word “commitment” — and mean four structurally different things by it.

Safety Cases: The Missing Methodology in Frontier AI Governance

What a safety case is as a methodology, GovAI’s rubric and safety-case papers, how they differ from SaferAI’s rubric, and the 2023 paper behind SB 53 and the EU AI Act.

Security Level: Nine Labs and Regulators, One Undefined Standard

Google DeepMind and Anthropic both cite RAND’s security-level scale but define different tiers; xAI, Meta, and the EU’s AI Act code reject grading and require a stated goal instead; and OpenAI’s own framework names a ‘Critical’ standard it says it hasn’t specified yet.

Who Counts as ‘Frontier AI’? Eleven Scope Tests Compared

The EU AI Act never says ‘frontier model.’ SB 53’s ‘catastrophic risk’ isn’t the EU’s ‘systemic risk.’ An 11-organization crosswalk of the if-then tests that decide who’s in scope — compute bars, revenue gates, relative tests, and one classified benchmark.

Capability Thresholds: 14 Labs and Regulators, One Undefined Term

Anthropic, OpenAI, and Google DeepMind all set “capability thresholds.” So do California, the EU, and the Frontier Model Forum. Almost none of them say what number triggers one — except Magic, whose AGI Readiness Policy discloses a single public number: 50% accuracy on LiveCodeBench.

The Statement on Superintelligence, Explained: Inside the Prohibition Call

What the Statement on Superintelligence actually says, who has verifiably signed it, and how its call for a development prohibition differs from a lab’s own model-specific threat-model vocabulary.

Explainable AI (XAI) Explained: Methods, Regulation, and Why It Matters

What explainable AI (XAI) means Explainable AI (XAI) is the property of an AI system, and the set of techniques used to achieve it, that lets a human understand why the system produced a particular output — which inputs mattered, how they were weighted, and what would have changed the result. IBM’s working definition frames […]

What Is AI Governance? A Complete Guide to Principles, Roles, and Compliance

AI governance is the policies, roles, and processes that keep AI systems safe, legal, and accountable. This guide defines the core principles, the roles involved, the major regulatory frameworks (EU AI Act, NIST AI RMF, ISO/IEC 42001, US state laws), and where compliance work actually happens.

NIKOLAI’s Track System: A Map of the Frontier AI Safety Landscape (N1-N10)

Not sure which frontier-AI-safety guide covers your question? NIKOLAI, CASRAI’s own dictionary of frontier-AI-safety elements, organizes the whole topic space into 10 tracks (N1-N10). Start here, pick the track closest to your question, and follow the links.

AI Terms, Acronyms, and Terminology: A Glossary for Governance and Compliance Teams

A plain-English glossary of common AI and machine-learning terms, plus AI governance acronyms decoded (RSP, FSF, NIST AI RMF, GPAI, CAISI, AISI) — for compliance and governance teams who aren’t yet fluent in AI vocabulary.

What Is Responsible AI? Principles, Frameworks, and How to Operationalize Them

Responsible AI principles — fairness, transparency, accountability, human oversight, security — explained, with how to turn them into a working framework and where CASRAI’s NIKOLAI dictionary fits in.

What Is AI Safety? A Plain-Language Guide

AI safety is the field working to prevent AI systems from causing harm, whether through misuse, accidents, or loss of control. Here’s the working vocabulary, how it differs from AI security, and who does this work.

What Is NIKOLAI? CASRAI’s Frontier-AI-Safety Dictionary Explained

NIKOLAI is CASRAI’s own frontier-AI-safety dictionary (nikolai-v0.2): 64 elements across 10 tracks covering thresholds, evaluations, safeguards, incidents, and accountability roles, plus Mapping Declarations, a public v1 REST API, and two MCP tools.

Frontier AI Labs: Who They Are and What Safety Frameworks They Publish

A directory of the developers generally treated as frontier AI labs under SB 53 and the EU AI Act, and which safety framework each one publishes, with links to CASRAI’s deep-dives on each.

The AI Safety Index: What It Rates and Who Publishes It

The Future of Life Institute publishes the AI Safety Index, grading frontier AI companies across six domains. Here’s what it rates and how it differs from SaferAI’s rubric.

“Concrete Problems in AI Safety”: The 2016 Paper Explained

What the 2016 paper “Concrete Problems in AI Safety” actually says: its five named problems — avoiding side effects, avoiding reward hacking, scalable supervision, safe exploration, and distributional shift — and why several of its authors later founded or joined Anthropic and OpenAI’s alignment teams.

The Center for AI Safety: What It Is and What It Does

An institutional profile of the Center for AI Safety (CAIS): its mission, the exact text of its 2023 Statement on AI Risk, its safety research and field-building work, and how it differs from Partnership on AI and the Frontier Model Forum.

What Is the AI Policy Institute (AIPI)?

An institutional profile of the AI Policy Institute (AIPI): a US 501(c)(3) that polls public opinion on AI and publishes policy research, founded by Daniel Colson.

Partnership on AI: What It Is and What It Does

Partnership on AI (PAI) is a nonprofit founded in 2016 that brings together industry, academic, and civil-society organisations to work on responsible AI development — distinct from industry-only bodies like the Frontier Model Forum.

AI Safety vs. AI Security: What the Distinction Actually Means

A short definitional guide distinguishing AI safety (preventing an AI system from causing unintended harm through its own behavior or capabilities) from AI security (protecting AI systems from external attack or misuse, such as model weight theft, adversarial attacks, and data poisoning), grounded in how frontier labs and oversight institutions actually use the terms.

AI Red Teaming: Definition and How It Works in Frontier AI Safety

AI red teaming inside a frontier safety framework is a safeguard-testing commitment, not the AppSec/jailbreak service most search results describe — here is what it actually tests, and how it differs from a dangerous-capability evaluation.

International AI Safety Report 2026: What It Found

What the Bengio-chaired International AI Safety Report 2026 found on frontier AI capabilities, malicious-use and control risks, and industry safety frameworks.

How to Grade a Frontier AI Safety Framework: The SaferAI Rubric

SaferAI, an independent nonprofit, scores frontier AI labs’ published safety frameworks against a four-dimension rubric. Here’s how the rubric works and how Anthropic, OpenAI, and Google DeepMind currently score.

What Is AI Alignment? Definition and Why It Matters for Frontier AI Safety

AI alignment is the technical problem of making a model behavior match its developers intended goals and values. This guide defines alignment, distinguishes it from AI safety and AI security, and shows how frontier labs evaluate it in practice.

CAISI: NIST’s Center for AI Standards and Innovation, Explained

CAISI is NIST’s Center for AI Standards and Innovation, renamed from the AI Safety Institute in June 2025. Its origin, its role inside NIST, its standards and evaluation mandate, and how it relates to the UK AI Security Institute.

AI Risk Assessment Framework and Risk Register: A Practical Starting Point

What a risk register actually contains (description, likelihood/impact, owner, mitigation, review cadence), how it maps to NIST’s AI RMF Govern-Map-Measure-Manage functions, and a practical starting template.

Responsible Scaling Policy (RSP): What It Is and How the Major Labs Compare

What a Responsible Scaling Policy is, how capability thresholds and safeguard tiers work, and how Anthropic’s RSP compares to OpenAI’s Preparedness Framework and Google DeepMind’s Frontier Safety Framework.

What Is a Frontier AI Model?

“Frontier AI model” is a capability class, not a specific product or lab. This guide walks through the three definitional approaches in active use — compute thresholds, capability thresholds, and relative state-of-the-art — and why the definition a framework uses is usually what triggers its obligations.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

AI policy question? Get an answer citing the framework.

An AI assistant specialized in research administration. Every answer links its sources to check before you act. 2 questions free, no account. $29/month after.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.