Concepts

Chinchilla scaling

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Chinchilla scaling allocates a fixed training-compute budget by growing model parameters and training tokens in roughly equal proportions.

1 · What it is

Earlier scaling work emphasized increasing parameter count and stopping large models before convergence. The Chinchilla study instead trained more than 400 models while varying both parameters and token counts. Its fitted optimum said that doubling model size should be accompanied by doubling training tokens.

Chinchilla tested that prediction with 70 billion parameters and roughly 1.4 trillion tokens. Gopher used 280 billion parameters and 300 billion training tokens. The two models used the same number of training FLOPs, yet Chinchilla outperformed Gopher and several larger models across a broad downstream evaluation set. Cerebras-GPT later reported a similar optimum of roughly 20 tokens per parameter on the Pile.

That ratio is a starting point for planning, not a law. The authors fitted it with runs that saw less than one pass over their data, and they suspect it may overstate the best size for the largest models. It also targets training compute only, not the cost of serving the model afterward. Once you count the cost of answering many requests, later work finds it pays to train a smaller model for longer than Chinchilla suggests.

2 · Why it exists

Parameter count alone can hide whether a language model received enough training data for its compute budget.

Oversized models are undertrainedA very large model trained on too few tokens can sit away from the lowest-loss allocation for its compute.
Training and serving differA smaller, longer-trained model can use the same pre-training compute while costing less to run.
The rule has a scopeThe original optimum minimizes pre-training loss for a fixed compute budget and leaves out the cost of serving the model.
3 · How it works

Hold training compute fixed, vary parameter count and token count, then choose the lowest-loss allocation.

Chinchilla scaling searches along a fixed-compute curve for the parameter-and-token allocation with minimum loss.
  1. 1 · fix computeChoose a pre-training FLOP budget and keep it constant across candidate runs.
  2. 2 · vary allocationPair larger models with fewer tokens and smaller models with more tokens while respecting that budget.
  3. 3 · compare lossFit the IsoFLOP loss curve and locate its minimum rather than choosing the largest parameter count.
  4. 4 · scale togetherFor larger budgets, increase model parameters and training tokens in approximately equal proportions.
  5. 5 · verifyTrain the chosen split, then compare its loss on evaluation text and its downstream task scores.

The optimum balances parameters and tokens for a stated compute budget.

4 · Where it's used
WhoWhat they askWhat it works with
Pre-training team“Should the next run use more parameters or more tokens?”IsoFLOP loss across candidate parameter-token pairs
Model platform“Can a smaller model match a larger model at the same training compute?”Training loss, task results, memory and inference cost
Budget owner“Which allocation gives the best measured quality per training FLOP?”The fitted compute-optimal frontier
Serving team“Does the training optimum match our lifetime cost optimum?”Expected request volume plus training and inference cost
5 · What it solves, and what it doesn't
solves
  • It turns a fixed pre-training budget into an empirical choice of model size and token count.
  • It identifies undertraining when parameter growth outruns data growth.
  • It explains how a smaller model can outperform a larger one trained with the same FLOPs.
  • It gives teams an experiment they can rerun, sweeping model size at several fixed budgets.
doesn't solve
  • The original rule does not include lifetime inference demand in its objective.
  • The fitted rule came from runs that saw less than one pass over their data. The authors also suspect it may overstate the best size for very large models.
  • A lower loss does not make a model safer. In the paper, toxic output barely tracked loss, and Chinchilla still showed bias.
  • More tokens only help if they are good tokens. The authors expect bigger datasets to pay off only when the text is high quality.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperTraining Compute-Optimal Large Language Models, DeepMind · read 27 Sept 2026
  2. officialAn empirical analysis of compute-optimal large language model training, Google DeepMind · read 27 Sept 2026
  3. paperScaling Laws for Neural Language Models, OpenAI · read 27 Sept 2026
  4. paperScaling Language Models: Methods, Analysis & Insights from Training Gopher, DeepMind · read 27 Sept 2026
  5. paperCerebras-GPT, Cerebras Systems · read 27 Sept 2026
  6. paperBeyond Chinchilla-Optimal: Accounting for Inference in Language Model Scaling Laws, Databricks MosaicML · read 27 Sept 2026