Jailbreaks
A jailbreak is a prompt written to trick an AI model into ignoring its safety training and producing something it would normally refuse.
A jailbreak is a trick prompt. Chatbots are trained to turn down harmful requests, and a jailbreak tries to talk the model out of that refusal. OWASP, a security standards group, files jailbreaking under prompt injection: input that makes the model drop its safety rules. In a jailbreak, the person typing is the attacker.
Most jailbreaks disguise the request. Some ask the model to play a character with no rules. Others paste in made-up chat turns where an “assistant” happily answers, or scramble the words with odd spelling or codes. Anthropic found that packing hundreds of fake example chats into one long prompt could override a model’s safety training. Researchers also built an automatic attack that searches for a string of odd characters to add to a request. That add-on makes a helpful-sounding reply more likely than a refusal. It worked on other public models too.
Defences usually come in layers. A small screening model, called a classifier, can check what a user types before the main model sees it. Another classifier can check the reply. Anthropic let outside testers try to break its prototype classifiers in a challenge called a bug bounty. They spent over 3,000 hours. No one found a single jailbreak that got answers to all ten banned test questions. Still, Anthropic knows of no model in real use that is fully safe from these attacks. So teams also watch the model’s answers for attacks that slip through.
Safety training teaches a model to say no, but that training has gaps attackers can find.
Follow one forbidden request through a guarded chatbot.
- 1 · askAsked plainly, a safety-trained model refuses a harmful request.
- 2 · wrapThe attacker disguises the request with role-play, fake dialogue turns or a character trick.
- 3 · screenAn input classifier can inspect the prompt before the main model sees it.
- 4 · answerIf the wrapper works, the model produces the harmful reply its training was meant to stop.
- 5 · checkAn output classifier can inspect the reply and block it.
A jailbreak changes the packaging of a request, not the request itself.
| Who | What they ask | What it works with |
|---|---|---|
| Chatbot developer | “Can users talk my support bot into breaking its rules?” | Incoming user prompts, screened before they reach the model |
| Safety researcher | “Which disguises still get past the model's refusals?” | Red-team prompts and the model's replies |
| Platform team | “Are jailbreak attempts getting through in production?” | Logged outputs, monitored for signs of a successful attack |
| Model provider | “Does a new defence block attacks without refusing normal questions?” | Refusal rates on harmless and harmful test prompts |
- Knowing common jailbreak styles lets tools flag role-play, fake dialogue and encoding attacks.
- Input screens can block a suspicious prompt before the main model answers.
- Input and output classifiers together stopped most jailbreaks in Anthropic's tests.
- Anthropic knows of no model in real use that is fully safe from jailbreaks.
- System prompts help, but OWASP says lasting protection needs ongoing updates to training and safety systems.
- Anthropic's first classifier prototype refused too many harmless questions and was costly to run.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialLLM01:2025 Prompt Injection, OWASP Gen AI Security Project · read 29 Sept 2026
- paperJailbroken: How Does LLM Safety Training Fail?, arXiv (Wei, Haghtalab and Steinhardt) · read 29 Sept 2026
- paperUniversal and Transferable Adversarial Attacks on Aligned Language Models, arXiv (Zou et al.) · read 29 Sept 2026
- officialMany-shot jailbreaking, Anthropic · read 29 Sept 2026
- officialConstitutional Classifiers: Defending against universal jailbreaks, Anthropic · read 29 Sept 2026
- docsPrompt Shields in Azure AI Content Safety, Microsoft Learn · read 29 Sept 2026
- docsMitigate jailbreaks and prompt injections, Claude Platform Docs · read 29 Sept 2026