Prompt injection
Prompt injection is text that sneaks new instructions into what an AI model reads, so the model behaves in unintended ways.
Prompt injection is text that sneaks new instructions into what an AI model reads, so the model does things nobody meant it to do. A “prompt” is the text you give an AI chatbot. “Injection” means pushing something in where it does not belong. The OWASP Gen AI Security Project, a group that tracks security risks, puts it first in its 2025 Top 10 risks for apps built on large language models (LLMs), the kind of AI behind chatbots.
The direct kind
In the direct kind, the person typing into the chat box tries to change how the model behaves. OWASP’s example is an attacker telling a support chatbot to ignore its rules, look up private data and send emails.
The sneaky kind
In the indirect kind, the attacker never talks to the model. They hide orders somewhere the app will probably read later. Researchers showed this works. Google lists emails, documents and calendar invites as hiding places.
Here is an everyday example. You ask an AI assistant to summarise a web page for your history homework. Somewhere on the page, perhaps in white text on a white background (our own illustration), a line tells the assistant to do something else. You never see it. The model reads it anyway. Orders can even hide inside pictures.
When the AI can act
Some AI tools, called agents, can take actions, not just chat. NIST, the US standards agency, calls this attack agent hijacking: someone else takes control. The poisoned data can look like a normal email, file or website. Anthropic warns that its Claude models may follow orders found in content, and advises keeping them away from sensitive data and risky actions.
One example, step by step
Follow the homework example through the figure. You ask for a summary. The assistant fetches the page, and the hidden line rides along. The app then joins its rules, your request and the page into one text. That is the key step: now the model cannot reliably tell whose words are whose. It may obey the hidden line and reach for its email tool. If the app asks you first, or never gave it that tool, the attack stops there.
What an attack can lead to
OWASP says an attack can leak private information, twist answers or steer important decisions. In NIST tests, hijacked agents could run an attacker’s code on the user’s computer or send all of a user’s cloud files to a stranger.
Not the same as a jailbreak
A jailbreak is one kind of prompt injection that makes the model drop its safety rules. Prompt injection is wider: many attacks just want the model to quietly do a different job.
Why it is hard to fix
Most language models handle instructions and data together, with no firm line between them. The hidden text only has to be read by the model, not seen by a person.
How people defend against it
- Least privilege. Give the model as little access as possible.
- Keep powerful jobs in normal code, not in the model’s hands.
- Check every tool use against what the user is allowed to do.
- Ask a human before risky actions.
- Label outside content as untrusted, which weakens its pull.
- Scan for attacks. Classifiers, programs that sort things into groups, can flag likely injections in what tools send back.
- Test it, treating the model like a stranger you do not trust.
The honest limits
None of these is a cure. No known method stops prompt injection completely. Retrieval (letting the model look things up) and fine-tuning (extra training on chosen examples) do not fully remove the risk. Telling the model to resist any rewrite of its rules helps, but guarantees nothing.
A language model reads its instructions and its data as one stream of text, and it cannot always tell them apart.
Follow one poisoned web page through an assistant that can send email.
- 1 · askThe user asks for something ordinary, such as a summary of a web page.
- 2 · fetchThe app pulls in outside content, and hidden instructions can ride along inside it.
- 3 · joinTrusted instructions and untrusted data are combined into one input for the model.
- 4 · actThe model may treat the hidden text as a command and try to use its tools.
- 5 · gateLimits on tools and a human approval step can stop the harmful action.
Prompt injection is a mixing problem: the model sees trusted orders and untrusted data as one block of text.
| Who | What they ask | What it works with |
|---|---|---|
| Student using an AI browser helper | “Can a web page I ask it to summarise make it do something I did not ask for” | Hidden instructions inside the fetched page |
| Company running a support chatbot | “Can a customer talk the bot into ignoring its rules and reading private records” | Direct injection typed into the chat box |
| Team building an email assistant | “What happens if an incoming email contains orders aimed at the assistant” | Indirect injection in the inbox the assistant reads |
| Security tester | “Do our trust boundaries hold when the model itself is treated as untrusted” | Attack simulations and penetration tests |
- Least-privilege access limits what a hijacked model can reach.
- Human approval for risky actions gives a person the final say.
- Marking outside content as untrusted reduces its pull on the model.
- Classifiers can flag likely injections in tool results such as screenshots.
- No known method prevents prompt injection completely.
- Retrieval and fine-tuning do not fully remove the risk.
- Asking the model to resist any effort to rewrite its rules is one mitigation, not a guarantee.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialLLM01:2025 Prompt Injection, OWASP Gen AI Security Project · read 28 Sept 2026
- docsLLM Prompt Injection Prevention Cheat Sheet, OWASP Cheat Sheet Series · read 28 Sept 2026
- officialTechnical Blog: Strengthening AI Agent Hijacking Evaluations, NIST Center for AI Standards and Innovation · read 28 Sept 2026
- paperNot what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, arXiv (Greshake et al.) · read 28 Sept 2026
- officialMitigating prompt injection attacks with a layered defense strategy, Google GenAI Security Team · read 28 Sept 2026
- docsComputer use tool, Anthropic · read 28 Sept 2026