FlashAttention
FlashAttention computes exact attention with tiling that reduces reads and writes between GPU memory levels.
FlashAttention is an IO-aware exact attention algorithm. It uses tiling to reduce reads and writes between GPU high-bandwidth memory and on-chip SRAM.
The algorithm splits Q, K and V into blocks. It loads blocks into SRAM, computes an attention block, updates the normalization factor, and writes the output back to HBM. It avoids materializing the large N by N attention matrix in HBM.
FlashAttention-2 improves work partitioning and parallelism. FlashAttention-3 targets Hopper GPUs. PyTorch notes that each fused scaled-dot-product attention kernel has specific input limitations.
Standard attention materializes large intermediate matrices in high-bandwidth memory.
Follow query, key and value blocks through tiled attention.
- 1 · tileSplit the query, key and value matrices into blocks.
- 2 · loadLoad blocks from high-bandwidth memory into fast on-chip SRAM.
- 3 · updateCompute attention block by block while maintaining softmax statistics.
- 4 · writeWrite the resulting output blocks back to high-bandwidth memory.
The algorithm changes the memory schedule while computing exact attention.
| Who | What they ask | What it works with |
|---|---|---|
| Model trainer | “Can exact attention use less GPU memory traffic?” | Query, key and value tiles |
| Kernel engineer | “Which blocks fit in on-chip SRAM?” | Tile sizes and memory capacity |
| Inference engineer | “Which scaled-dot-product attention kernel will run?” | Backend selection and input constraints |
- FlashAttention uses tiling to reduce memory reads and writes.
- The first paper reported faster training than existing baselines.
- FlashAttention-2 improved work partitioning and parallelism.
- FlashAttention still computes exact attention rather than replacing the attention rule.
- Kernel availability depends on hardware and input limitations.
- FlashAttention-3 targets Hopper GPUs.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperFlashAttention, Dao et al. · read 28 Sept 2026
- paperFlashAttention-2, Dao · read 28 Sept 2026
- paperFlashAttention-3, Shah et al. · read 28 Sept 2026
- repoflash-attention, Dao-AILab · read 28 Sept 2026
- docsscaled_dot_product_attention, PyTorch · read 28 Sept 2026