Concepts

Jailbreaks

4 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A jailbreak is a prompt written to trick an AI model into ignoring its safety training and producing something it would normally refuse.

1 · What it is

A jailbreak is a trick prompt. Chatbots are trained to turn down harmful requests, and a jailbreak tries to talk the model out of that refusal. OWASP, a security standards group, files jailbreaking under prompt injection: input that makes the model drop its safety rules. In a jailbreak, the person typing is the attacker.

Most jailbreaks disguise the request. Some ask the model to play a character with no rules. Others paste in made-up chat turns where an “assistant” happily answers, or scramble the words with odd spelling or codes. Anthropic found that packing hundreds of fake example chats into one long prompt could override a model’s safety training. Researchers also built an automatic attack that searches for a string of odd characters to add to a request. That add-on makes a helpful-sounding reply more likely than a refusal. It worked on other public models too.

Defences usually come in layers. A small screening model, called a classifier, can check what a user types before the main model sees it. Another classifier can check the reply. Anthropic let outside testers try to break its prototype classifiers in a challenge called a bug bounty. They spent over 3,000 hours. No one found a single jailbreak that got answers to all ten banned test questions. Still, Anthropic knows of no model in real use that is fully safe from these attacks. So teams also watch the model’s answers for attacks that slip through.

2 · Why it exists

Safety training teaches a model to say no, but that training has gaps attackers can find.

Goals can clashResearchers found that a model's wish to be helpful can pull against its safety goals, and attackers exploit that tension.
Training misses cornersSafety training may not carry over to every kind of input the model can still understand, such as unusual encodings.
Long inputs help attackersBigger context windows let an attacker pack in hundreds of fake example chats that push the model toward answering.
3 · How it works

Follow one forbidden request through a guarded chatbot.

The jailbreak lives in the wrapper around the request. Defences sit before and after the model.
  1. 1 · askAsked plainly, a safety-trained model refuses a harmful request.
  2. 2 · wrapThe attacker disguises the request with role-play, fake dialogue turns or a character trick.
  3. 3 · screenAn input classifier can inspect the prompt before the main model sees it.
  4. 4 · answerIf the wrapper works, the model produces the harmful reply its training was meant to stop.
  5. 5 · checkAn output classifier can inspect the reply and block it.

A jailbreak changes the packaging of a request, not the request itself.

4 · Where it's used
WhoWhat they askWhat it works with
Chatbot developer“Can users talk my support bot into breaking its rules?”Incoming user prompts, screened before they reach the model
Safety researcher“Which disguises still get past the model's refusals?”Red-team prompts and the model's replies
Platform team“Are jailbreak attempts getting through in production?”Logged outputs, monitored for signs of a successful attack
Model provider“Does a new defence block attacks without refusing normal questions?”Refusal rates on harmless and harmful test prompts
5 · What it solves, and what it doesn't
solves
  • Knowing common jailbreak styles lets tools flag role-play, fake dialogue and encoding attacks.
  • Input screens can block a suspicious prompt before the main model answers.
  • Input and output classifiers together stopped most jailbreaks in Anthropic's tests.
doesn't solve
  • Anthropic knows of no model in real use that is fully safe from jailbreaks.
  • System prompts help, but OWASP says lasting protection needs ongoing updates to training and safety systems.
  • Anthropic's first classifier prototype refused too many harmless questions and was costly to run.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialLLM01:2025 Prompt Injection, OWASP Gen AI Security Project · read 29 Sept 2026
  2. paperJailbroken: How Does LLM Safety Training Fail?, arXiv (Wei, Haghtalab and Steinhardt) · read 29 Sept 2026
  3. paperUniversal and Transferable Adversarial Attacks on Aligned Language Models, arXiv (Zou et al.) · read 29 Sept 2026
  4. officialMany-shot jailbreaking, Anthropic · read 29 Sept 2026
  5. officialConstitutional Classifiers: Defending against universal jailbreaks, Anthropic · read 29 Sept 2026
  6. docsPrompt Shields in Azure AI Content Safety, Microsoft Learn · read 29 Sept 2026
  7. docsMitigate jailbreaks and prompt injections, Claude Platform Docs · read 29 Sept 2026