Grouped-query attention
Grouped-query attention lets each group of query heads share one key head and one value head.
Loading cached keys and values can bottleneck autoregressive decoding. Grouped-query attention uses an intermediate number of key-value heads.
It keeps multiple query heads and divides them into groups. Each query group shares one key head and one value head. One group gives multi-query attention; one group per query head gives multi-head attention.
Multi-query attention shares keys and values across all attention heads. Llama 2 uses GQA to improve inference scalability for its larger models. Llama 3 uses eight key-value heads to improve inference speed and reduce key-value cache size during decoding.
Group count remains an architecture choice. Converting an existing multi-head checkpoint requires constructing group key and value heads.
Loading cached keys and values can bottleneck autoregressive decoding.
Group query heads around shared key-value heads.
- 1 · query headsKeep multiple query heads.
- 2 · groupsDivide those query heads into groups.
- 3 · shareGive each group one shared key head and value head.
- 4 · attendLet every query head attend through its group key and value heads.
One group gives multi-query attention; one group per query head gives multi-head attention.
| Who | What they ask | What it works with |
|---|---|---|
| Uptrained Transformer | “Can an MHA checkpoint be converted to fewer key-value heads?” | Mean-pooled key and value heads within each group |
| Llama 2 70B | “How can inference scalability improve in a larger model?” | Grouped-query attention |
| Llama 3 | “How can decoding cache size and inference speed improve?” | Eight key-value heads shared by more query heads |
- It reduces the number of key-value heads relative to multi-head attention.
- It provides an intermediate configuration between multi-head and multi-query attention.
- The original GQA study reported quality close to multi-head attention with speed comparable to multi-query attention.
- Group count remains an architecture choice.
- Converting an existing multi-head checkpoint requires constructing group key and value heads.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperGQA Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints, Ainslie et al. · read 28 Sept 2026
- paperFast Transformer Decoding One Write-Head is All You Need, Shazeer · read 28 Sept 2026
- paperLlama 2 Open Foundation and Fine-Tuned Chat Models, Touvron et al. · read 28 Sept 2026
- docsLlama model documentation, Hugging Face · read 28 Sept 2026
- paperThe Llama 3 Herd of Models, Dubey et al. · read 28 Sept 2026