KV cacheConcepts

Key-value cache

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

A KV cache lets a language model save work from earlier tokens, so each new token does not redo the same maths.

1 · What it is

A language model reads and writes text in small pieces called tokens. Inside the model, attention turns every token into three lists of numbers: a query, a key and a value. Once a token has been handled, these do not change for later tokens, so the stored pairs can be reused. So there is no reason to work them out twice. A KV cache stores them and hands them back when they are needed.

Picture a model that has already written 999 tokens of a story. To write token 1,000, it needs information from all 999 earlier tokens. Token 1,001 needs that same information again, plus token 1,000. With a cache, the model looks the old keys and values up instead of recomputing them, and only works out the pair for the newest token.

The catch is memory. Each layer keeps its own cache, and a basic cache gets longer with every token. Multi-query attention shrinks the cache by letting all the attention heads share one set of keys and values. Grouped-query attention sits in the middle: a few groups of heads each share one set. Serving systems, the software that runs a model for many users, also manage the space carefully. Badly managed cache memory is wasted, and that limits how many requests fit at once. PagedAttention, used by vLLM, borrows the idea of memory pages from operating systems to cut that waste.

2 · Why it exists

Language models write one token at a time, and every new token depends on all the tokens before it.

Repeated workWithout a cache, the model recomputes the keys and values of every earlier token at each step.
Memory useThe cache can fill a large share of a chip's memory, which becomes a limit for long texts.
Slow readingLoading all the saved keys and values again and again can make each step slow.
3 · How it works

Follow one new token through one attention layer.

Only the newest token's key and value are computed. The older ones come straight from the cache.
  1. 1 · storeThe model keeps a separate cache of keys and values in each attention layer.
  2. 2 · computeFor the newest token only, it computes one new key and one new value.
  3. 3 · appendIt adds that key and value to the end of the cache.
  4. 4 · attendAttention combines the new pair with all the saved pairs to work out its scores.
4 · Where it's used
WhoWhat they askWhat it works with
Chat app developer“Why do long conversations use more GPU memory?”The KV cache kept for each conversation
Serving engineer“How many requests fit on one GPU at once?”KV cache memory per request
Model designer“Can the cache be made smaller?”Multi-query or grouped-query attention
Student running a model at home“Why does generation stop with an out-of-memory error?”Moving the cache to the computer's main memory
5 · What it solves, and what it doesn't
solves
  • It stops the model from recomputing keys and values for earlier tokens.
  • It cuts computation time, so responses come back faster.
  • Serving systems such as vLLM can share cache memory within and across requests.
doesn't solve
  • It does not let the model write several tokens at once; it still predicts one at a time.
  • It does not keep memory fixed, because a basic cache grows with every new token.
  • It does not stop the model from loading all the earlier keys and values at each step.
  • It is meant for running a model, not for training one.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsHow caching works, Hugging Face Transformers · read 28 Sept 2026
  2. docsCache strategies, Hugging Face Transformers · read 28 Sept 2026
  3. paperFast Transformer Decoding: One Write-Head is All You Need, arXiv · read 28 Sept 2026
  4. paperGQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints, arXiv · read 28 Sept 2026
  5. paperEfficient Memory Management for Large Language Model Serving with PagedAttention, arXiv · read 28 Sept 2026