Concepts

Music generation models

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

A music generation model can generate music in the raw-audio domain.

1 · What it is

Jukebox generates music in the raw-audio domain. It compresses raw audio into discrete codes and models those codes with autoregressive Transformers.

MusicLM treats conditional music generation as a hierarchical sequence-to-sequence task. MusicGen instead uses a single-stage Transformer language model over streams of compressed discrete music representations.

MusicGen notes that representing audio as several parallel token streams comes with a modeling cost. An inexact parallel decomposition can accumulate errors over time. MusicLM can misunderstand negations and does not adhere to precise temporal ordering described in text.

AudioLM demonstrates coherent piano continuations without symbolic music representations.

2 · Why it exists

Music generation can require long-term structure.

Compress audioJukebox uses a multi-scale VQ-VAE to compress raw audio into discrete codes.
Keep structureMusic Transformer targets musical pieces with long-term structure.
Follow a promptMusicLM casts text-conditioned music generation as hierarchical sequence-to-sequence modeling.
3 · How it works

Follow one text prompt through MusicLM's hierarchical sequence task.

MusicLM describes text-conditioned music generation as a hierarchical sequence-to-sequence task.
  1. 1 · conditionUse a text description to condition music generation.
  2. 2 · planModel conditional generation as a hierarchical sequence-to-sequence task.
  3. 3 · detailUse a hierarchy of sequences in that generation task.
  4. 4 · renderGenerate music from the text description.

MusicGen is a single-stage Transformer language model over compressed discrete music representations.

4 · Where it's used
WhoWhat they askWhat it works with
Music tool builder“Does the generated clip follow the requested description?”Text and generated audio pairs
Model researcher“How does this system represent music during generation?”The model's sequence representation
Runtime engineer“Which stage dominates generation time?”Token generation and audio decoding timings
5 · What it solves, and what it doesn't
solves
  • Text can condition music generation.
  • Discrete codes can reduce the sequence length of raw audio.
  • Autoregressive Transformers can model compressed music codes.
  • AudioLM can generate coherent piano continuations without symbolic music representations.
doesn't solve
  • Modeling several parallel streams of discrete audio tokens has a cost.
  • Inexact parallel prediction can accumulate errors over time.
  • MusicLM can misunderstand negations in a text description.
  • MusicLM does not adhere to precise temporal ordering described in text.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperMusicLM: Generating Music From Text, Agostinelli and colleagues · read 28 Sept 2026
  2. paperSimple and Controllable Music Generation, Copet and colleagues · read 28 Sept 2026
  3. paperJukebox: A Generative Model for Music, Dhariwal and colleagues · read 28 Sept 2026
  4. paperAudioLM: a Language Modeling Approach to Audio Generation, Borsos and colleagues · read 28 Sept 2026
  5. paperMusic Transformer, Huang and colleagues · read 28 Sept 2026