Music generation models
A music generation model can generate music in the raw-audio domain.
Jukebox generates music in the raw-audio domain. It compresses raw audio into discrete codes and models those codes with autoregressive Transformers.
MusicLM treats conditional music generation as a hierarchical sequence-to-sequence task. MusicGen instead uses a single-stage Transformer language model over streams of compressed discrete music representations.
MusicGen notes that representing audio as several parallel token streams comes with a modeling cost. An inexact parallel decomposition can accumulate errors over time. MusicLM can misunderstand negations and does not adhere to precise temporal ordering described in text.
AudioLM demonstrates coherent piano continuations without symbolic music representations.
Music generation can require long-term structure.
Follow one text prompt through MusicLM's hierarchical sequence task.
- 1 · conditionUse a text description to condition music generation.
- 2 · planModel conditional generation as a hierarchical sequence-to-sequence task.
- 3 · detailUse a hierarchy of sequences in that generation task.
- 4 · renderGenerate music from the text description.
MusicGen is a single-stage Transformer language model over compressed discrete music representations.
| Who | What they ask | What it works with |
|---|---|---|
| Music tool builder | “Does the generated clip follow the requested description?” | Text and generated audio pairs |
| Model researcher | “How does this system represent music during generation?” | The model's sequence representation |
| Runtime engineer | “Which stage dominates generation time?” | Token generation and audio decoding timings |
- Text can condition music generation.
- Discrete codes can reduce the sequence length of raw audio.
- Autoregressive Transformers can model compressed music codes.
- AudioLM can generate coherent piano continuations without symbolic music representations.
- Modeling several parallel streams of discrete audio tokens has a cost.
- Inexact parallel prediction can accumulate errors over time.
- MusicLM can misunderstand negations in a text description.
- MusicLM does not adhere to precise temporal ordering described in text.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperMusicLM: Generating Music From Text, Agostinelli and colleagues · read 28 Sept 2026
- paperSimple and Controllable Music Generation, Copet and colleagues · read 28 Sept 2026
- paperJukebox: A Generative Model for Music, Dhariwal and colleagues · read 28 Sept 2026
- paperAudioLM: a Language Modeling Approach to Audio Generation, Borsos and colleagues · read 28 Sept 2026
- paperMusic Transformer, Huang and colleagues · read 28 Sept 2026