State space models
A continuous-time linear state space model uses x-dot = Ax + Bu for state evolution and y = Cx + Du for readout.
In continuous time, the state derivative is x-dot = Ax + Bu and the readout is y = Cx + Du. After discretization, a recurrent form updates the state from the previous state and current input.
LSSL uses a trainable subset of structured state matrices. S4 conditions A with a low-rank correction so it can be diagonalized stably and reduced to a Cauchy-kernel computation. S4D is a diagonal version of S4 whose paper reports matching S4 results.
S5 uses one multi-input, multi-output system and a parallel scan. HiPPO compresses signals by projecting them onto polynomial bases.
Conventional sequence models can struggle to scale to very long sequences.
Follow one input through a discrete state update.
- 1 · receiveThe layer receives the current input and the previous state.
- 2 · updateMatrices A and B combine that information into the next state.
- 3 · readMatrices C and D map the state and input to the current output.
- 4 · repeatThe new state passes to the next position in the sequence.
S4 introduced a new state-space parameterization.
| Who | What they ask | What it works with |
|---|---|---|
| Audio researcher | “How can a model carry information across a long waveform?” | A state space layer |
| Time-series engineer | “Can one layer map an input signal to an output signal?” | A linear state update and readout |
| Sequence-model researcher | “Can recurrent computation also be parallelized for training?” | Convolutional or parallel-scan form |
- S4 reports that its state space model can be computed more efficiently than prior approaches.
- S5 uses a parallel scan while matching the computational efficiency of S4.
- A generic state matrix can have prohibitive computation and memory requirements.
- S4D reports that initialization is critical for performance.
- The Mamba paper identifies inability to perform content-based reasoning as a weakness.
- Diagonal state space models still involve choices in parameterization and computation.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperCombining Recurrent, Convolutional, and Continuous-time Models with Linear State-Space Layers, Gu et al. · read 28 Sept 2026
- paperEfficiently Modeling Long Sequences with Structured State Spaces, Gu, Goel and Ré · read 28 Sept 2026
- paperOn the Parameterization and Initialization of Diagonal State Space Models, Gu et al. · read 28 Sept 2026
- paperSimplified State Space Layers for Sequence Modeling, Smith, Warrington and Linderman · read 28 Sept 2026
- paperHiPPO: Recurrent Memory with Optimal Polynomial Projections, Gu et al. · read 28 Sept 2026
- paperMamba: Linear-Time Sequence Modeling with Selective State Spaces, Gu and Dao · read 28 Sept 2026