Chinchilla scaling
Chinchilla scaling allocates a fixed training-compute budget by growing model parameters and training tokens in roughly equal proportions.
Earlier scaling work emphasized increasing parameter count and stopping large models before convergence. The Chinchilla study instead trained more than 400 models while varying both parameters and token counts. Its fitted optimum said that doubling model size should be accompanied by doubling training tokens.
Chinchilla tested that prediction with 70 billion parameters and roughly 1.4 trillion tokens. Gopher used 280 billion parameters and 300 billion training tokens. The two models used the same number of training FLOPs, yet Chinchilla outperformed Gopher and several larger models across a broad downstream evaluation set. Cerebras-GPT later reported a similar optimum of roughly 20 tokens per parameter on the Pile.
That ratio is a starting point for planning, not a law. The authors fitted it with runs that saw less than one pass over their data, and they suspect it may overstate the best size for the largest models. It also targets training compute only, not the cost of serving the model afterward. Once you count the cost of answering many requests, later work finds it pays to train a smaller model for longer than Chinchilla suggests.
Parameter count alone can hide whether a language model received enough training data for its compute budget.
Hold training compute fixed, vary parameter count and token count, then choose the lowest-loss allocation.
- 1 · fix computeChoose a pre-training FLOP budget and keep it constant across candidate runs.
- 2 · vary allocationPair larger models with fewer tokens and smaller models with more tokens while respecting that budget.
- 3 · compare lossFit the IsoFLOP loss curve and locate its minimum rather than choosing the largest parameter count.
- 4 · scale togetherFor larger budgets, increase model parameters and training tokens in approximately equal proportions.
- 5 · verifyTrain the chosen split, then compare its loss on evaluation text and its downstream task scores.
The optimum balances parameters and tokens for a stated compute budget.
| Who | What they ask | What it works with |
|---|---|---|
| Pre-training team | “Should the next run use more parameters or more tokens?” | IsoFLOP loss across candidate parameter-token pairs |
| Model platform | “Can a smaller model match a larger model at the same training compute?” | Training loss, task results, memory and inference cost |
| Budget owner | “Which allocation gives the best measured quality per training FLOP?” | The fitted compute-optimal frontier |
| Serving team | “Does the training optimum match our lifetime cost optimum?” | Expected request volume plus training and inference cost |
- It turns a fixed pre-training budget into an empirical choice of model size and token count.
- It identifies undertraining when parameter growth outruns data growth.
- It explains how a smaller model can outperform a larger one trained with the same FLOPs.
- It gives teams an experiment they can rerun, sweeping model size at several fixed budgets.
- The original rule does not include lifetime inference demand in its objective.
- The fitted rule came from runs that saw less than one pass over their data. The authors also suspect it may overstate the best size for very large models.
- A lower loss does not make a model safer. In the paper, toxic output barely tracked loss, and Chinchilla still showed bias.
- More tokens only help if they are good tokens. The authors expect bigger datasets to pay off only when the text is high quality.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperTraining Compute-Optimal Large Language Models, DeepMind · read 27 Sept 2026
- officialAn empirical analysis of compute-optimal large language model training, Google DeepMind · read 27 Sept 2026
- paperScaling Laws for Neural Language Models, OpenAI · read 27 Sept 2026
- paperScaling Language Models: Methods, Analysis & Insights from Training Gopher, DeepMind · read 27 Sept 2026
- paperCerebras-GPT, Cerebras Systems · read 27 Sept 2026
- paperBeyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws, Databricks MosaicML · read 27 Sept 2026