Concepts

Mamba

4 min readadvancedUpdated 28 Sept 2026
1 · In one line

Mamba is a sequence architecture whose state space parameters depend on the input, letting each token control what information is kept or forgotten.

1 · What it is

Mamba is a sequence architecture built from selective state space layers. Its state space parameters are functions of the current input. A token can therefore change what information is propagated or forgotten.

That selection prevents the use of efficient convolutions. Mamba instead uses a hardware-aware parallel algorithm in recurrent mode. The resulting architecture has linear sequence-length scaling. The original Mamba design does not use attention or MLP blocks.

The official repository is based on a selective state space layer. The official repository provides repeating Mamba blocks and a language-model head. Mamba-2’s core layer is a refinement of Mamba’s selective state space model. MambaByte is a token-free Mamba adaptation trained autoregressively on byte sequences. Jamba interleaves Mamba with Transformer layers.

2 · Why it exists

State space models can struggle with content-based reasoning.

Content mattersThe model can propagate or forget information depending on the current token.
Long sequencesStandard Transformer computation becomes expensive as sequence length grows.
Hardware pathInput-dependent state parameters prevent the use of efficient convolutions.
3 · How it works

Follow tokens through selection and the recurrent scan.

Each token sets state space parameters. The selective scan runs in recurrent mode.
  1. 1 · receiveThe selective state space layer receives the input sequence.
  2. 2 · selectFunctions of the current input produce token-dependent state space parameters.
  3. 3 · scanA hardware-aware parallel algorithm runs in recurrent mode.
  4. 4 · returnThe official implementation returns an output tensor with the same shape as its input.

Mamba's selection is input-dependent. The original architecture does not use attention.

4 · Where it's used
WhoWhat they askWhat it works with
Language-model researcher“Can a recurrent state react to the current token?”A selective state space layer
Systems engineer“How does sequence-length cost grow in this architecture?”The hardware-aware selective scan
Byte-model researcher“Can Mamba process autoregressive byte sequences without tokens?”MambaByte's token-free byte-sequence setup
5 · What it solves, and what it doesn't
solves
  • Input-dependent parameters let the model selectively propagate or forget information according to the current token.
  • The original architecture scales linearly with sequence length.
  • The official implementation includes an example of a complete language model.
doesn't solve
  • Input-dependent selection prevents the use of efficient convolutions.
  • GPU execution for the official selective-scan extension has hardware and software requirements.
  • The original Mamba block contains neither attention nor MLP blocks.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperMamba: Linear-Time Sequence Modeling with Selective State Spaces, Gu and Dao · read 28 Sept 2026
  2. repoMamba official implementation, State Spaces · read 28 Sept 2026
  3. paperTransformers are SSMs: Generalized Models and Efficient Algorithms Through Structured State Space Duality, Dao and Gu · read 28 Sept 2026
  4. paperMambaByte: Token-free Selective State Space Model, Wang et al. · read 28 Sept 2026
  5. paperJamba: A Hybrid Transformer-Mamba Language Model, Lieber et al. · read 28 Sept 2026