Prompt caching
Prompt caching saves the work a model did on the unchanged start of a prompt, so the next request with the same start is faster and cheaper.
When a model reads a prompt, it turns every token into internal numbers. These numbers are called key-value states. Think of them as the model’s notes on what it has read. Many apps send the same long start with every request: the same tools, the same instructions, the same document. Without caching, the model redoes the work on that start every time.
Prompt caching keeps those states for the unchanged start of a prompt, called the prefix. When a later request begins with exactly the same prefix, the saved states are reused. Only the new part, such as the latest question, is processed fresh. Picture a support bot whose long policy text sits at the top of every chat. After the first customer, the policy text is reused at a lower price, and only the new question is processed from scratch.
The match must be exact, from the very first token. If anything changes early on, everything after that change is processed again. So builders put fixed content first and changing content last. The cache stores the model’s notes, not an answer. So replies are the same as without caching.
The big providers offer it. OpenAI turns prompt caching on by default for supported models. Google switches on implicit caching automatically for its Gemini models from version 2.5 onward. Anthropic lets developers mark where the cached part ends with a field called cache_control. Saved states expire after a while. On Claude the default is five minutes, and the timer resets each time they are used. Caches are kept separate between organizations.
For engineers: On OpenAI’s models from GPT-5.6 onward, writing a prefix costs 1.25 times the normal input rate and each reuse costs 0.1 times. So ten calls that share a prefix (the first saves it, the other nine reuse it) cost 2.15 times the normal input price, not 10 times. The vLLM server calls this automatic prefix caching and reuses the saved states of an earlier query that shares the same prefix. A 2023 research paper called Prompt Cache reused these saved states across prompts. It cut the wait for the first word by 8 times on GPUs (graphics chips) and up to 60 times on CPUs (ordinary processors). The SGLang system reuses saved states with a method called RadixAttention.
Many requests to a model start with the same long text, and the model reprocesses it every time.
Follow two requests that share the same start.
- 1 · orderThe app puts the parts that never change, like tools and instructions, at the start of the prompt.
- 2 · writeOn the first request, the model processes that start and saves its internal states in a cache.
- 3 · matchA later request is checked for a start that exactly matches something already saved.
- 4 · reuseOn a match, the saved states are reused and only the new part is processed.
Prompt caching reuses the model's work on the prompt, not a stored answer.
| Who | What they ask | What it works with |
|---|---|---|
| Customer support bot | “Where is my order?” | The same long policy and tool list at the start of every chat |
| Coding assistant | “Now fix the failing test” | The growing conversation history from earlier turns |
| Research team | “What does section four of this report say about costs?” | One long document asked about many times |
| Agent developer | “Which tool should run next?” | The fixed tool definitions and system instructions |
- The start of the prompt is not recomputed on each request.
- Reused input tokens are billed at a lower cached rate.
- The answer can start sooner because less input needs processing.
- The answer is the same as it would be without caching.
- It does not speed up writing a long answer.
- It does not help when the start of the prompt keeps changing.
- Saved states expire after a while, so a slow follow-up may miss the cache.
- Short prompts below a minimum length are not cached.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsPrompt caching, Anthropic · read 28 Sept 2026
- docsPrompt caching, OpenAI · read 28 Sept 2026
- docsContext caching, Google AI for Developers · read 28 Sept 2026
- docsAutomatic Prefix Caching, vLLM · read 28 Sept 2026
- paperPrompt Cache: Modular Attention Reuse for Low-Latency Inference, Gim et al., MLSys 2024 · read 28 Sept 2026
- paperSGLang: Efficient Execution of Structured Language Model Programs, Zheng et al. · read 28 Sept 2026