Skip to main content
v2026.11,858 entries · CC-BY 4.0

What Anthropic Pays for a Jailbreak: Inside Its Model Safety Bug Bounty

Anthropic will pay up to $15,000 through HackerOne for a novel, universal jailbreak against its CBRN and cybersecurity safeguards. Here’s exactly what that program covers, what it doesn’t, and where it sits inside Anthropic’s own safety commitments.

Written and maintained by CASRAI Editorial Board

Last updated

Last verified: September 20, 2026, against Anthropic’s own announcement at anthropic.com/news/model-safety-bug-bounty (published August 8, 2024). CASRAI’s guide to AI red teaming mentions this program exactly once, in a single quoted phrase from Anthropic’s Responsible Scaling Policy — “bug bounty program” — and moves on. That guide is upfront about why it stays abstract: its own FAQ notes that “OpenAI’s Preparedness Framework and Google DeepMind’s Frontier Safety Framework… were not directly accessible while researching this guide,” so it compares red-teaming commitments at the level of policy language rather than program detail. This page goes the other direction. It stays inside the one program with a public, dated, primary-source announcement — Anthropic’s Model Safety Bug Bounty — and gets concrete: what it actually paid, on what platform, against what scope, and to whom.

What Anthropic’s Model Safety Bug Bounty Actually Pays For

On August 8, 2024, Anthropic announced a bug bounty program run in partnership with HackerOne, offering rewards up to $15,000 for novel, universal jailbreak attacks. The program did not cover jailbreaks in general. It targeted “critical, high risk domains such as CBRN (chemical, biological, radiological, and nuclear) and cybersecurity” — the two categories Anthropic treats as most consequential if a safeguard fails. Anthropic defined a universal jailbreak narrowly, too: “a type of vulnerability in AI systems that allows a user to consistently bypass the safety measures across a wide range of topics.” A jailbreak that worked once, on one prompt, in one narrow context was not the target; a jailbreak that reliably defeated the safeguard across many different harmful requests was.

That specificity is the whole point of this page. A dollar figure, a named platform, and a two-domain scope are the kind of detail that turns “we run a bug bounty program” from a line in a policy document into something an outside researcher, journalist, or compliance reader can actually evaluate.

What Was Actually Being Tested

The program wasn’t testing Anthropic’s models as shipped. It was testing an unreleased, next-generation safety mitigation system — safeguards Anthropic had built but not yet deployed publicly. That framing matters: this was pre-deployment adversarial testing of a specific safeguard layer, closer in spirit to the “external red team and penetration testing experts” language in Anthropic’s Responsible Scaling Policy than to an open-ended vulnerability-disclosure program covering the entire product surface.

Who Could Apply, and When

Access was invite-only at launch. Anthropic asked researchers to submit an application by Friday, August 16, 2024 — an eight-day window from the announcement — and said it would follow up with selected applicants in the fall of 2024. Anthropic’s stated target for that first cohort was “experienced AI security researchers” with demonstrated expertise “identifying jailbreaks in language models,” not a general-public bounty open to anyone with a HackerOne account. The announcement also described plans to expand the program more broadly over time, though this page verifies only the invite-only launch terms as published on August 8, 2024, not any later expansion — a later stage of the program, if one exists, is outside what was checked for this piece.

Where This Sits Inside Anthropic’s Broader Safety Commitments

The bug bounty isn’t a standalone program; Anthropic’s own Responsible Scaling Policy folds it into a larger loop. Per the RSP text already surfaced in CASRAI’s AI red-teaming guide, Anthropic keeps its real-time misuse classifiers current using “insights from our incident response protocol, bug bounty program, and internal and external red-teaming efforts.” Read against that sentence, the Model Safety Bug Bounty is one of at least three named inputs — alongside incident response and red-teaming — that feed the same deployment safeguards. A universal jailbreak found and paid out through HackerOne doesn’t just earn a researcher $15,000; on Anthropic’s own account, it becomes training signal for the classifiers protecting the next release.

One limit worth stating plainly: this page does not make any claim about equivalent programs at OpenAI or Google DeepMind. Their Bugcrowd- and VRP-style disclosure programs are the natural comparison, and both were checked for this piece — both attempts returned 403 and 404 responses rather than an accessible, citable program page. That’s a gap in what could be verified this session, not a claim that no such program exists. Anthropic’s Model Safety Bug Bounty is the one program with a live, dated, primary-source announcement backing every figure on this page, so it’s the only one this page describes in dollar terms.

The NIKOLAI Angle: Does This Fit an Existing Element?

NIKOLAI is CASRAI’s own independent, unendorsed dictionary of frontier-AI-safety elements — not a standard any lab has adopted, and not something Anthropic has reviewed or confirmed. Three of its elements looked like plausible fits for a HackerOne-run bounty program, and it’s worth being honest about how well each one actually holds up under that program’s specifics.

Evaluator access attestation doesn’t fit. That element is built around negotiated, documented access — NIKOLAI’s page describes Anthropic offering “employee-like access” to embedded evaluators, with unredacted risk reports and system access under agreed confidentiality terms. A HackerOne researcher testing a jailbreak through a bounty platform is the opposite arrangement: no negotiated access grant, no embedded relationship, just a scoped target and a payout schedule.

AI-model review doesn’t fit either, for a different reason — it’s specifically about an AI system performing assurance work (an AI model reviewing documents, monitoring agent actions, running automated red-teaming), not a human researcher testing a model from outside.

External review comes closest. NIKOLAI defines it as “a documented assessment conducted by a party independent of the model developer, evaluating a model, risk report, safeguard, or framework compliance” — and a HackerOne researcher probing an unreleased safeguard is, on the plain words of that definition, exactly that: an independent party evaluating a safeguard. But NIKOLAI’s own crosswalk for this element, built from eleven organizations’ published material, is populated by mandatory third-party-evaluator arrangements — Anthropic’s Long-Term Benefit Trust approving reviewer selection, the EU GPAI Code’s “adequately qualified independent external evaluators,” METR’s formal assessment relationships. None of those rows describes a bounty-platform researcher paid per finding. CASRAI checked for that distinction specifically and didn’t find it: NIKOLAI’s crosswalk, as published, does not explicitly separate bounty-hunter-style external testing from the negotiated third-party-evaluator relationships it does document. That’s a genuine gap in NIKOLAI’s own coverage, not a confirmed tie-in — the same kind of honest limitation the red-teaming guide flags about its own OpenAI and Google DeepMind sourcing. Readers who want NIKOLAI’s fuller treatment of this territory should see the N8 — Transparency and review track page, where External review lives alongside Evaluator access attestation and AI-model review.

Frequently Asked Questions

How much does Anthropic pay for a jailbreak?

Up to $15,000, for a novel, universal jailbreak against the CBRN and cybersecurity safeguards in scope for Anthropic’s Model Safety Bug Bounty, run on HackerOne. Anthropic defined “universal” as a vulnerability that consistently bypasses safety measures across a wide range of topics, not a one-off bypass of a single prompt.

Is this program still accepting applications?

The invite-only application window Anthropic announced closed on August 16, 2024, eight days after the program’s August 8, 2024 launch. This page verifies only that launch announcement; it does not verify whether Anthropic has since opened a broader or successor version of the program, and readers checking for current eligibility should confirm directly with Anthropic or HackerOne.

What was actually being tested?

An unreleased, next-generation safety mitigation system — safeguards Anthropic had built but not yet shipped publicly — rather than Anthropic’s already-deployed models.

Does Anthropic’s bug bounty program have a NIKOLAI element?

Not a confirmed one. NIKOLAI’s External review element is the closest match by definition, but its published crosswalk doesn’t explicitly cover bounty-platform testers as distinct from negotiated third-party evaluators — see the NIKOLAI section above for what was checked and what wasn’t found.

Do OpenAI and Google DeepMind run comparable programs?

Almost certainly something exists at both, but this page doesn’t assert program details for either. Attempts to reach OpenAI’s Bugcrowd-hosted program and Google’s AI Vulnerability Reward Program both returned inaccessible responses (403 and 404) during research for this piece, so neither is described here in the concrete terms this page uses for Anthropic’s program.

Related Reading

Follow CASRAI

Research-administration guidance, standards updates and independent tool reviews.

Ask CASRAI · free to try

Ask about What Anthropic Pays for a Jailbreak: Inside Its Model Safety Bug Bounty

Ask your first 2 questions free below. Subscribers get 150 a day for $29 a month.

Ask CASRAI answers research-administration questions and cites the passages behind every claim. When our sources don't cover a question, it says so.

Answers draw on CASRAI's guides and dictionary plus the federal and funder documents we index: Federal Register, Grants.gov, Regulations.gov and UKRI.

Works on this site and inside Claude, Cursor and the AI tools you already use.

Everything CASRAI publishes — this page, the dictionary, the guides and the news — stays free to read, with no account and no card.

Referenced across the research world

University of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logoUniversity of Cambridge logoColumbia University logoCrossref logoUniversity of Edinburgh logoHarvard University logoUniversity of Oxford logoPrinceton University logoStanford School of Medicine logoUniversity College London logoORCID logo
  • University of Cambridge logo
  • Columbia University logo
  • Crossref logo
  • University of Edinburgh logo
  • Harvard University logo
  • University of Oxford logo
  • Princeton University logo
  • Stanford School of Medicine logo
  • University College London logo
  • ORCID logo

View CASRAI adoption →