Concepts

FlashAttention

3 min readadvancedUpdated 28 Sept 2026
1 · In one line

FlashAttention computes exact attention with tiling that reduces reads and writes between GPU memory levels.

1 · What it is

FlashAttention is an IO-aware exact attention algorithm. It uses tiling to reduce reads and writes between GPU high-bandwidth memory and on-chip SRAM.

The algorithm splits Q, K and V into blocks. It loads blocks into SRAM, computes an attention block, updates the normalization factor, and writes the output back to HBM. It avoids materializing the large N by N attention matrix in HBM.

FlashAttention-2 improves work partitioning and parallelism. FlashAttention-3 targets Hopper GPUs. PyTorch notes that each fused scaled-dot-product attention kernel has specific input limitations.

2 · Why it exists

Standard attention materializes large intermediate matrices in high-bandwidth memory.

Memory trafficFlashAttention is designed around reads and writes between levels of GPU memory.
Large matrixStandard attention materializes the N by N attention matrix to high-bandwidth memory.
Exact resultFlashAttention is an exact attention algorithm, not an attention approximation.
3 · How it works

Follow query, key and value blocks through tiled attention.

FlashAttention tiled memory schedule Blocks of query, key and value matrices move from high-bandwidth memory to on-chip SRAM. A tiled exact-attention step updates softmax statistics and output blocks before the output returns to high-bandwidth memory.
Tiling keeps attention blocks in fast memory and avoids materializing the complete score matrix in high-bandwidth memory.
  1. 1 · tileSplit the query, key and value matrices into blocks.
  2. 2 · loadLoad blocks from high-bandwidth memory into fast on-chip SRAM.
  3. 3 · updateCompute attention block by block while maintaining softmax statistics.
  4. 4 · writeWrite the resulting output blocks back to high-bandwidth memory.

The algorithm changes the memory schedule while computing exact attention.

4 · Where it's used
WhoWhat they askWhat it works with
Model trainer“Can exact attention use less GPU memory traffic?”Query, key and value tiles
Kernel engineer“Which blocks fit in on-chip SRAM?”Tile sizes and memory capacity
Inference engineer“Which scaled-dot-product attention kernel will run?”Backend selection and input constraints
5 · What it solves, and what it doesn't
solves
  • FlashAttention uses tiling to reduce memory reads and writes.
  • The first paper reported faster training than existing baselines.
  • FlashAttention-2 improved work partitioning and parallelism.
doesn't solve
  • FlashAttention still computes exact attention rather than replacing the attention rule.
  • Kernel availability depends on hardware and input limitations.
  • FlashAttention-3 targets Hopper GPUs.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperFlashAttention, Dao et al. · read 28 Sept 2026
  2. paperFlashAttention-2, Dao · read 28 Sept 2026
  3. paperFlashAttention-3, Shah et al. · read 28 Sept 2026
  4. repoflash-attention, Dao-AILab · read 28 Sept 2026
  5. docsscaled_dot_product_attention, PyTorch · read 28 Sept 2026