Red teaming
Red teaming means deliberately testing an AI system the way an attacker would, to find its weak spots.
Red teaming borrows its name from computer security. There, a red team is a group given permission to act like an attacker. The defenders it tests are called the blue team. With large language models, the phrase has grown to mean almost any way of poking at, testing or attacking the system.
The people matter as much as the prompts. Good red teams mix people with an attacker’s mindset and ordinary users who did not build the product. Security experts can probe a model for jailbreaks, which are prompts that trick it past its rules, and for extraction of its hidden instructions. Medical experts can help find risks in a chatbot built for health care providers.
People are slow, so machines help too. Researchers can use one language model to generate test cases for another. Anthropic has used a model to generate attacks and then trained a model on the results to resist similar attacks. One research team shared 38,961 of its recorded attacks publicly so other people could learn from them. Microsoft offers PyRIT, a free, openly published toolkit that helps security teams spot risks in generative AI systems.
Language models can harm users in ways that are hard to predict.
Follow one round of red teaming from plan to fix.
- 1 · planAssign testers to specific harms or product features.
- 2 · probeTesters explore openly to uncover as many kinds of harm as they can.
- 3 · recordEach finding is logged with the input prompt and the system's output.
- 4 · listFindings become a list of harms that guides testers in later rounds.
- 5 · fixThe team tests versions with and without fixes to see whether the fixes work.
Red teaming finds problems, but it does not replace systematic measurement.
| Who | What they ask | What it works with |
|---|---|---|
| Chatbot safety team | “Can our assistant be talked into giving dangerous advice?” | Replies to prompts written to trick it |
| Health app builder | “What could go wrong when doctors use our chatbot?” | Test conversations run by medical experts |
| AI lab researcher | “Can one model find prompts that break another?” | Test cases written by a language model |
| App developer | “Do the default filters still hold inside our app?” | The base model and the app, before and after fixes |
- It aims to find and fix bad model behaviour before it reaches users.
- It checks whether default safety filters leave gaps in a specific app.
- It produces a list of harms that tells a team what to measure and fix.
- In one study, model-written tests uncovered tens of thousands of offensive chatbot replies.
- A few striking examples do not show how common a harm is.
- Teams use different methods, so their results are hard to compare.
- Coverage can be lopsided; Anthropic says most of its red teaming takes place in English.
- Finding a problem does not fix it; fixes must be built and retested.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialred team - Glossary, NIST Computer Security Resource Center · read 28 Sept 2026
- officialChallenges in red teaming AI systems, Anthropic · read 28 Sept 2026
- docsPlanning red teaming for large language models (LLMs) and their applications, Microsoft Learn · read 28 Sept 2026
- paperRed Teaming Language Models with Language Models, arXiv · read 28 Sept 2026
- paperRed Teaming Language Models to Reduce Harms: Methods, Scaling Behaviors, and Lessons Learned, arXiv · read 28 Sept 2026
- repoPython Risk Identification Tool for generative AI (PyRIT), Microsoft · read 28 Sept 2026