Sparsity
Sparsity means that many possible model values or computation paths are zero, absent, or inactive for a given input.
Sparsity describes missing work, and it comes from several places. Pruning leaves zeros in the weight matrices. During inference, the common ReLU function also makes many activations exactly zero. A mixture-of-experts model goes further: it sends each input through only a few experts and leaves the rest switched off.
The pattern decides what the hardware can do with it. Zeros scattered anywhere give the most freedom, but the storage format has to keep each nonzero’s position. NVIDIA’s 2:4 pattern allows at most two nonzeros in every group of four, which gives regular memory access and even work across cores. The trade is accuracy: tighter patterns usually cost a little more of it. Experts skip most of the model per token, but tokens can crowd onto one expert. That expert then overflows and drops them, so Switch Transformer training adds a loss term that spreads tokens out.
So report two things: how many zeros there are, and what they actually saved. A sparse format can even run slower than the dense one. Measure memory, time, power, and accuracy on the hardware you will really use.
A dense model runs every weight for every input, even when many of the numbers are zero.
Match the sparse pattern to an execution strategy.
- 1 · createProduce zero weights, zero activations, or experts that stay off for a given input.
- 2 · patternChoose unstructured zeros, a fixed N:M pattern, or expert routing.
- 3 · encodeStore the nonzeros with their positions, or the routing choice, in a layout the runtime supports.
- 4 · executeRun a sparse kernel or the chosen expert, then time it on the real hardware.
A model can have many parameters yet low per-token compute when it uses conditional activation, as in mixture-of-experts routing.
| Who | What they ask | What it works with |
|---|---|---|
| Hardware engineer | “Which pattern maps to this accelerator?” | Supported 2:4 formats and speedup by matrix size |
| Model researcher | “Should topology stay fixed during training?” | Dynamic sparse training results |
| Systems operator | “Is expert routing balanced across devices?” | Expert overflow and load-balancing loss |
- Sparse encodings can reduce storage and data movement.
- Supported patterns can skip zero-valued operations.
- Conditional sparsity can increase parameter capacity without activating every parameter per input.
- A sparse format can run slower than the dense version.
- Position metadata, routing, and communication add overhead.
- Pruning too hard hurts quality, and uneven routing can overload some experts.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperRigging the Lottery: Making All Tickets Winners, Evci et al. · read 27 Sept 2026
- paperSwitch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity, Fedus, Zoph, and Shazeer · read 27 Sept 2026
- paperSCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks, Parashar et al. · read 27 Sept 2026
- officialExploiting NVIDIA Ampere Structured Sparsity with cuSPARSELt, NVIDIA · read 27 Sept 2026
- paperSparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Frantar and Alistarh · read 27 Sept 2026
- docstorch.sparse, PyTorch · read 27 Sept 2026