Concepts

Grouped-query attention

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Grouped-query attention lets each group of query heads share one key head and one value head.

1 · What it is

Loading cached keys and values can bottleneck autoregressive decoding. Grouped-query attention uses an intermediate number of key-value heads.

It keeps multiple query heads and divides them into groups. Each query group shares one key head and one value head. One group gives multi-query attention; one group per query head gives multi-head attention.

Multi-query attention shares keys and values across all attention heads. Llama 2 uses GQA to improve inference scalability for its larger models. Llama 3 uses eight key-value heads to improve inference speed and reduce key-value cache size during decoding.

Group count remains an architecture choice. Converting an existing multi-head checkpoint requires constructing group key and value heads.

2 · Why it exists

Loading cached keys and values can bottleneck autoregressive decoding.

Cache trafficDecoder inference repeatedly loads attention keys and values.
Two endpointsMulti-head attention has separate key-value heads, while multi-query attention shares one pair across all query heads.
Middle settingGrouped-query attention uses an intermediate number of key-value heads.
3 · How it works

Group query heads around shared key-value heads.

Each query group shares one key head and one value head.
  1. 1 · query headsKeep multiple query heads.
  2. 2 · groupsDivide those query heads into groups.
  3. 3 · shareGive each group one shared key head and value head.
  4. 4 · attendLet every query head attend through its group key and value heads.

One group gives multi-query attention; one group per query head gives multi-head attention.

4 · Where it's used
WhoWhat they askWhat it works with
Uptrained Transformer“Can an MHA checkpoint be converted to fewer key-value heads?”Mean-pooled key and value heads within each group
Llama 2 70B“How can inference scalability improve in a larger model?”Grouped-query attention
Llama 3“How can decoding cache size and inference speed improve?”Eight key-value heads shared by more query heads
5 · What it solves, and what it doesn't
solves
  • It reduces the number of key-value heads relative to multi-head attention.
  • It provides an intermediate configuration between multi-head and multi-query attention.
  • The original GQA study reported quality close to multi-head attention with speed comparable to multi-query attention.
doesn't solve
  • Group count remains an architecture choice.
  • Converting an existing multi-head checkpoint requires constructing group key and value heads.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperGQA Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints, Ainslie et al. · read 28 Sept 2026
  2. paperFast Transformer Decoding One Write-Head is All You Need, Shazeer · read 28 Sept 2026
  3. paperLlama 2 Open Foundation and Fine-Tuned Chat Models, Touvron et al. · read 28 Sept 2026
  4. docsLlama model documentation, Hugging Face · read 28 Sept 2026
  5. paperThe Llama 3 Herd of Models, Dubey et al. · read 28 Sept 2026