Concepts

State space models

4 min readadvancedUpdated 28 Sept 2026
1 · In one line

A continuous-time linear state space model uses x-dot = Ax + Bu for state evolution and y = Cx + Du for readout.

1 · What it is

In continuous time, the state derivative is x-dot = Ax + Bu and the readout is y = Cx + Du. After discretization, a recurrent form updates the state from the previous state and current input.

LSSL uses a trainable subset of structured state matrices. S4 conditions A with a low-rank correction so it can be diagonalized stably and reduced to a Cauchy-kernel computation. S4D is a diagonal version of S4 whose paper reports matching S4 results.

S5 uses one multi-input, multi-output system and a parallel scan. HiPPO compresses signals by projecting them onto polynomial bases.

2 · Why it exists

Conventional sequence models can struggle to scale to very long sequences.

Long dependenciesConventional sequence models can struggle to scale to sequences of ten thousand or more steps.
Efficient computationS4 reports more efficient computation than prior approaches.
3 · How it works

Follow one input through a discrete state update.

The recurrent view updates state step by step. S5 uses a parallel scan.
  1. 1 · receiveThe layer receives the current input and the previous state.
  2. 2 · updateMatrices A and B combine that information into the next state.
  3. 3 · readMatrices C and D map the state and input to the current output.
  4. 4 · repeatThe new state passes to the next position in the sequence.

S4 introduced a new state-space parameterization.

4 · Where it's used
WhoWhat they askWhat it works with
Audio researcher“How can a model carry information across a long waveform?”A state space layer
Time-series engineer“Can one layer map an input signal to an output signal?”A linear state update and readout
Sequence-model researcher“Can recurrent computation also be parallelized for training?”Convolutional or parallel-scan form
5 · What it solves, and what it doesn't
solves
  • S4 reports that its state space model can be computed more efficiently than prior approaches.
  • S5 uses a parallel scan while matching the computational efficiency of S4.
doesn't solve
  • A generic state matrix can have prohibitive computation and memory requirements.
  • S4D reports that initialization is critical for performance.
  • The Mamba paper identifies inability to perform content-based reasoning as a weakness.
  • Diagonal state space models still involve choices in parameterization and computation.