Concepts

Prompt injection

5 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Prompt injection is text that sneaks new instructions into what an AI model reads, so the model behaves in unintended ways.

1 · What it is

Prompt injection is text that sneaks new instructions into what an AI model reads, so the model does things nobody meant it to do. A “prompt” is the text you give an AI chatbot. “Injection” means pushing something in where it does not belong. The OWASP Gen AI Security Project, a group that tracks security risks, puts it first in its 2025 Top 10 risks for apps built on large language models (LLMs), the kind of AI behind chatbots.

The direct kind

In the direct kind, the person typing into the chat box tries to change how the model behaves. OWASP’s example is an attacker telling a support chatbot to ignore its rules, look up private data and send emails.

The sneaky kind

In the indirect kind, the attacker never talks to the model. They hide orders somewhere the app will probably read later. Researchers showed this works. Google lists emails, documents and calendar invites as hiding places.

Here is an everyday example. You ask an AI assistant to summarise a web page for your history homework. Somewhere on the page, perhaps in white text on a white background (our own illustration), a line tells the assistant to do something else. You never see it. The model reads it anyway. Orders can even hide inside pictures.

When the AI can act

Some AI tools, called agents, can take actions, not just chat. NIST, the US standards agency, calls this attack agent hijacking: someone else takes control. The poisoned data can look like a normal email, file or website. Anthropic warns that its Claude models may follow orders found in content, and advises keeping them away from sensitive data and risky actions.

One example, step by step

Follow the homework example through the figure. You ask for a summary. The assistant fetches the page, and the hidden line rides along. The app then joins its rules, your request and the page into one text. That is the key step: now the model cannot reliably tell whose words are whose. It may obey the hidden line and reach for its email tool. If the app asks you first, or never gave it that tool, the attack stops there.

What an attack can lead to

OWASP says an attack can leak private information, twist answers or steer important decisions. In NIST tests, hijacked agents could run an attacker’s code on the user’s computer or send all of a user’s cloud files to a stranger.

Not the same as a jailbreak

A jailbreak is one kind of prompt injection that makes the model drop its safety rules. Prompt injection is wider: many attacks just want the model to quietly do a different job.

Why it is hard to fix

Most language models handle instructions and data together, with no firm line between them. The hidden text only has to be read by the model, not seen by a person.

How people defend against it

  • Least privilege. Give the model as little access as possible.
  • Keep powerful jobs in normal code, not in the model’s hands.
  • Check every tool use against what the user is allowed to do.
  • Ask a human before risky actions.
  • Label outside content as untrusted, which weakens its pull.
  • Scan for attacks. Classifiers, programs that sort things into groups, can flag likely injections in what tools send back.
  • Test it, treating the model like a stranger you do not trust.

The honest limits

None of these is a cure. No known method stops prompt injection completely. Retrieval (letting the model look things up) and fine-tuning (extra training on chosen examples) do not fully remove the risk. Telling the model to resist any rewrite of its rules helps, but guarantees nothing.

2 · Why it exists

A language model reads its instructions and its data as one stream of text, and it cannot always tell them apart.

No clear boundaryMost language models process instructions and data together, with no firm line between the two.
Hidden from peopleInjected text does not have to be visible to a human; it only has to be read by the model.
Tools raise the stakesWhen a model can call functions or run commands, a successful injection can reach those tools too.
3 · How it works

Follow one poisoned web page through an assistant that can send email.

The attack works at the join, where trusted instructions and untrusted content become one prompt. Defences sit after it, between the model and its tools.
  1. 1 · askThe user asks for something ordinary, such as a summary of a web page.
  2. 2 · fetchThe app pulls in outside content, and hidden instructions can ride along inside it.
  3. 3 · joinTrusted instructions and untrusted data are combined into one input for the model.
  4. 4 · actThe model may treat the hidden text as a command and try to use its tools.
  5. 5 · gateLimits on tools and a human approval step can stop the harmful action.

Prompt injection is a mixing problem: the model sees trusted orders and untrusted data as one block of text.

4 · Where it's used
WhoWhat they askWhat it works with
Student using an AI browser helper“Can a web page I ask it to summarise make it do something I did not ask for”Hidden instructions inside the fetched page
Company running a support chatbot“Can a customer talk the bot into ignoring its rules and reading private records”Direct injection typed into the chat box
Team building an email assistant“What happens if an incoming email contains orders aimed at the assistant”Indirect injection in the inbox the assistant reads
Security tester“Do our trust boundaries hold when the model itself is treated as untrusted”Attack simulations and penetration tests
5 · What it solves, and what it doesn't
solves
  • Least-privilege access limits what a hijacked model can reach.
  • Human approval for risky actions gives a person the final say.
  • Marking outside content as untrusted reduces its pull on the model.
  • Classifiers can flag likely injections in tool results such as screenshots.
doesn't solve
  • No known method prevents prompt injection completely.
  • Retrieval and fine-tuning do not fully remove the risk.
  • Asking the model to resist any effort to rewrite its rules is one mitigation, not a guarantee.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. officialLLM01:2025 Prompt Injection, OWASP Gen AI Security Project · read 28 Sept 2026
  2. docsLLM Prompt Injection Prevention Cheat Sheet, OWASP Cheat Sheet Series · read 28 Sept 2026
  3. officialTechnical Blog: Strengthening AI Agent Hijacking Evaluations, NIST Center for AI Standards and Innovation · read 28 Sept 2026
  4. paperNot what you've signed up for: Compromising Real-World LLM-Integrated Applications with Indirect Prompt Injection, arXiv (Greshake et al.) · read 28 Sept 2026
  5. officialMitigating prompt injection attacks with a layered defense strategy, Google GenAI Security Team · read 28 Sept 2026
  6. docsComputer use tool, Anthropic · read 28 Sept 2026