Skip to main content
v2026.11,858 entries · CC-BY 4.0
Dictionary termTrack CStablev2026.2

Jailbreak (LLM)

A prompt or interaction pattern that causes a language model to bypass its safety training and produce outputs the model was tuned to refuse, such as harmful instructions, restricted content, or violations of provider policy.

ByCASRAI Editorial Board
· Last updated 5 Sept 2026
Share this

Ask CASRAI · free to try

Ask about Jailbreak (LLM)

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

An AI assistant specialized in research administration. It cites the sources behind every answer, labels web answers and says when it can't answer.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Examples

Worked examples

  • Is an instance

    A role-play prompt ('pretend you are an unrestricted AI...') that elicits otherwise-refused content.

  • Is an instance

    An adversarial-suffix attack appending a learned token sequence that disables refusal.

Counter-examples

Looks similar, but isn't

  • Not an instance

    A legitimate user query in scope of the model's intended use.

  • Not an instance

    A prompt-injection attack delivered via tool input (different category).

Editorial commentary

Jailbreaks include role-play framings, multi-turn manipulation, encoding tricks (base64, ROT13), and adversarial-suffix attacks (Zou et al., 2023). Resistance to jailbreaking is a target of post-training (RLHF, constitutional AI) and a focus of red-team evaluation. The category is a moving target: each newly disclosed jailbreak typically prompts new mitigations.

References

  • Zou et al., 'Universal and Transferable Adversarial Attacks on Aligned Language Models' (arXiv 2023); Wei, Haghtalab, Steinhardt, 'Jailbroken: How Does LLM Safety Training Fail?' (NeurIPS 2023).

Frequently Asked Questions

How does an LLM jailbreak actually work?

Jailbreaks use techniques such as role-play framings, multi-turn manipulation, encoding tricks like base64 or ROT13, and adversarial-suffix attacks that append a crafted token sequence to a prompt. Each technique tries to get the model to produce output its safety training was meant to refuse.

Is an LLM jailbreak a security exploit?

No. A jailbreak is an alignment failure, not a software exploit: it works by crafting a prompt that causes the model to bypass its safety training, rather than by exploiting a bug in code.

What is the difference between a jailbreak and a prompt injection attack?

A jailbreak is a prompt or interaction pattern aimed directly at the model to bypass its own refusal behavior. Prompt injection is a related but distinct category, where the attack is delivered through tool input rather than the user’s own prompt. See Prompt injection.

How do AI providers try to prevent jailbreaks?

Resistance to jailbreaking is a target of post-training methods such as RLHF (reinforcement learning from human feedback) and constitutional AI, and it is a specific focus of red-team evaluation both before and after a model is released.

Can jailbreaks be fixed permanently?

Not in a lasting way. Jailbreak resistance is a moving target: each newly disclosed jailbreak technique typically prompts new mitigations, rather than closing the issue once and for all.

Also known as

LLM jailbreak

Machine-readable encodings

Use in your systems

JATS XML <role> element
xml
<role vocab="credit"
      vocab-identifier="https://casrai.org/dictionary/"
      vocab-term="Jailbreak (LLM)"
      vocab-term-identifier="https://casrai.org/dictionary/term/jailbreak-llm" />
Schema.org DefinedTerm (JSON-LD)
json
{
  "@context": "https://schema.org",
  "@type": "DefinedTerm",
  "@id": "https://casrai.org/dictionary/term/jailbreak-llm",
  "name": "Jailbreak (LLM)",
  "identifier": "https://casrai.org/dictionary/term/jailbreak-llm",
  "description": "A prompt or interaction pattern that causes a language model to bypass its safety training and produce outputs the model was tuned to refuse, such as harmful instructions, restricted content, or violations of provider policy.",
  "inDefinedTermSet": "https://casrai.org/dictionary/domain/ai-ml-research-outputs#set",
  "url": "https://casrai.org/dictionary/term/jailbreak-llm",
  "alternateName": [
    "LLM jailbreak"
  ],
  "license": "https://creativecommons.org/licenses/by/4.0/",
  "publisher": {
    "@id": "https://casrai.org/#organization"
  },
  "author": {
    "@id": "https://casrai.org/#editorial-team"
  },
  "datePublished": "2026-05-21T02:22:51",
  "dateModified": "2026-09-05T14:24:52",
  "inLanguage": "en-GB",
  "isAccessibleForFree": true
}

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →

Ask CASRAI · Regulatory Radar

AI policy question? Get an answer citing the framework.

An AI assistant for research administration. Ask about 2 CFR 200, NIH, NSF or IRB. 2 questions free, no account. $29/month after, cancel anytime.

  • Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.
  • Every answer numbers its sources and links each one, so you can check the source yourself.