Mamba
Mamba is a sequence architecture whose state space parameters depend on the input, letting each token control what information is kept or forgotten.
Mamba is a sequence architecture built from selective state space layers. Its state space parameters are functions of the current input. A token can therefore change what information is propagated or forgotten.
That selection prevents the use of efficient convolutions. Mamba instead uses a hardware-aware parallel algorithm in recurrent mode. The resulting architecture has linear sequence-length scaling. The original Mamba design does not use attention or MLP blocks.
The official repository is based on a selective state space layer. The official repository provides repeating Mamba blocks and a language-model head. Mamba-2’s core layer is a refinement of Mamba’s selective state space model. MambaByte is a token-free Mamba adaptation trained autoregressively on byte sequences. Jamba interleaves Mamba with Transformer layers.
State space models can struggle with content-based reasoning.
Follow tokens through selection and the recurrent scan.
- 1 · receiveThe selective state space layer receives the input sequence.
- 2 · selectFunctions of the current input produce token-dependent state space parameters.
- 3 · scanA hardware-aware parallel algorithm runs in recurrent mode.
- 4 · returnThe official implementation returns an output tensor with the same shape as its input.
Mamba's selection is input-dependent. The original architecture does not use attention.
| Who | What they ask | What it works with |
|---|---|---|
| Language-model researcher | “Can a recurrent state react to the current token?” | A selective state space layer |
| Systems engineer | “How does sequence-length cost grow in this architecture?” | The hardware-aware selective scan |
| Byte-model researcher | “Can Mamba process autoregressive byte sequences without tokens?” | MambaByte's token-free byte-sequence setup |
- Input-dependent parameters let the model selectively propagate or forget information according to the current token.
- The original architecture scales linearly with sequence length.
- The official implementation includes an example of a complete language model.
- Input-dependent selection prevents the use of efficient convolutions.
- GPU execution for the official selective-scan extension has hardware and software requirements.
- The original Mamba block contains neither attention nor MLP blocks.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMamba: Linear-Time Sequence Modeling with Selective State Spaces, Gu and Dao · read 28 Sept 2026
- repoMamba official implementation, State Spaces · read 28 Sept 2026
- paperTransformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, Dao and Gu · read 28 Sept 2026
- paperMambaByte: Token-free Selective State Space Model, Wang et al. · read 28 Sept 2026
- paperJamba: A Hybrid Transformer-Mamba Language Model, Lieber et al. · read 28 Sept 2026