Concepts

Sparsity

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Sparsity means that many possible model values or computation paths are zero, absent, or inactive for a given input.

1 · What it is

Sparsity describes missing work, and it comes from several places. Pruning leaves zeros in the weight matrices. During inference, the common ReLU function also makes many activations exactly zero. A mixture-of-experts model goes further: it sends each input through only a few experts and leaves the rest switched off.

The pattern decides what the hardware can do with it. Zeros scattered anywhere give the most freedom, but the storage format has to keep each nonzero’s position. NVIDIA’s 2:4 pattern allows at most two nonzeros in every group of four, which gives regular memory access and even work across cores. The trade is accuracy: tighter patterns usually cost a little more of it. Experts skip most of the model per token, but tokens can crowd onto one expert. That expert then overflows and drops them, so Switch Transformer training adds a loss term that spreads tokens out.

So report two things: how many zeros there are, and what they actually saved. A sparse format can even run slower than the dense one. Measure memory, time, power, and accuracy on the hardware you will really use.

2 · Why it exists

A dense model runs every weight for every input, even when many of the numbers are zero.

Wasted multiplyMultiplying by zero adds nothing to the answer, yet dense hardware still does it.
Wasted movementStoring and moving zeros uses memory and bandwidth.
IrregularityWhen zeros sit anywhere, the format must record where each nonzero lives.
3 · How it works

Match the sparse pattern to an execution strategy.

Different sparsity patterns require different execution strategiesUnstructured zeros, regular two-of-four zeros, and one-of-four expert routing are encoded into supported layouts, which a matching runtime then executes and measures.Unstructured weights× 0 × 0 0 × 0 ×arbitrary positionsStructured 2:4× 0 × 0 | 0 × 0 ×regular groupsExpert routingE1 E2 [E3] E4one active pathKEY STEPEncode the patternindices for arbitrary zerosfixed layout for 2:4token routes for expertsin a layout the runtime readsMatching runtimesparse kernel or routermeasure utilizationThe same sparsity percentage can have very different runtime costs.
Sparsity is a property; acceleration requires a representation and kernel designed for its pattern.
  1. 1 · createProduce zero weights, zero activations, or experts that stay off for a given input.
  2. 2 · patternChoose unstructured zeros, a fixed N:M pattern, or expert routing.
  3. 3 · encodeStore the nonzeros with their positions, or the routing choice, in a layout the runtime supports.
  4. 4 · executeRun a sparse kernel or the chosen expert, then time it on the real hardware.

A model can have many parameters yet low per-token compute when it uses conditional activation, as in mixture-of-experts routing.

4 · Where it's used
WhoWhat they askWhat it works with
Hardware engineer“Which pattern maps to this accelerator?”Supported 2:4 formats and speedup by matrix size
Model researcher“Should topology stay fixed during training?”Dynamic sparse training results
Systems operator“Is expert routing balanced across devices?”Expert overflow and load-balancing loss
5 · What it solves, and what it doesn't
solves
  • Sparse encodings can reduce storage and data movement.
  • Supported patterns can skip zero-valued operations.
  • Conditional sparsity can increase parameter capacity without activating every parameter per input.
doesn't solve
  • A sparse format can run slower than the dense version.
  • Position metadata, routing, and communication add overhead.
  • Pruning too hard hurts quality, and uneven routing can overload some experts.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperRigging the Lottery: Making All Tickets Winners, Evci et al. · read 27 Sept 2026
  2. paperSwitch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity, Fedus, Zoph, and Shazeer · read 27 Sept 2026
  3. paperSCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks, Parashar et al. · read 27 Sept 2026
  4. officialExploiting NVIDIA Ampere Structured Sparsity with cuSPARSELt, NVIDIA · read 27 Sept 2026
  5. paperSparseGPT: Massive Language Models Can Be Accurately Pruned in One-Shot, Frantar and Alistarh · read 27 Sept 2026
  6. docstorch.sparse, PyTorch · read 27 Sept 2026