Emergent abilities
Emergent abilities are task capabilities that appear absent in smaller models but present at a larger scale.
The original definition focuses on the shape of measured task performance across scale. BIG-bench evaluated models spanning millions to hundreds of billions of parameters, and its authors saw three patterns on different tasks, steady gains, no gains, and sudden breakthroughs. The sudden jumps often came on tasks where the model must chain several parts together, or where the scoring rule is fragile.
But a jump in a chart is not automatically a jump in the underlying behavior. The metric critique showed that a nonlinear or discontinuous score can turn smooth output changes into an apparent cliff. With continuous metrics, the same fixed outputs can instead form a smooth and predictable curve.
The debate did not end there. Later loss-based work reported thresholds even under continuous metrics, and found that models with different parameter and data sizes can perform similarly at the same pre-training loss under controlled conditions. In the UL2R experiments, a changed training recipe made scores on a few BIG-bench tasks jump at 62B parameters, whereas the original PaLM only jumped at 540B. So emergence is still an open debate, not a settled law of scaling.
A sudden benchmark jump can change what a model seems able to do and how teams forecast it.
Test the same task across model scales, then inspect the measurement itself.
- 1 · hold fixedPick one task and one model family, and save the outputs so every metric scores exactly the same answers.
- 2 · sweepEvaluate several model sizes rather than comparing only one small and one large model.
- 3 · inspectPlot both the task's headline metric and a continuous measure of partial progress when one is meaningful.
- 4 · retestCheck other tasks and other model families before treating one threshold as a general rule.
The key question is whether the behavior jumps or only the reported score does.
| Who | What they ask | What it works with |
|---|---|---|
| Evaluation team | “Did reasoning truly appear at this scale?” | Per-example outputs under exact and continuous metrics |
| Model researcher | “Which checkpoint first exceeds chance?” | Task score against model scale and pre-training loss |
| Safety team | “Could a risky capability appear between releases?” | Capability sweeps across intermediate checkpoints |
| Product lead | “Will a larger model handle this workflow?” | Repeated task evaluations under the intended prompt and metric |
- The concept names sharp task-level changes that simple small-model extrapolation can miss.
- Scale sweeps can reveal where performance first moves above random guessing.
- Metric audits separate a true behavioral discontinuity from a thresholding artefact.
- Loss-based analysis offers another axis for comparing models with different parameter and data sizes.
- A benchmark jump does not reveal the internal mechanism that produced it.
- One sharp result does not prove that all models or prompts will show the same threshold.
- A cliff seen under one metric may flatten into a smooth curve under another.
- Parameter count alone does not predict task scores, since models of different sizes trained on different amounts of data can score alike.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperEmergent Abilities of Large Language Models, Google Research · read 27 Sept 2026
- officialCharacterizing Emergent Phenomena in Large Language Models, Google Research · read 27 Sept 2026
- paperAre Emergent Abilities of Large Language Models a Mirage?, NeurIPS · read 27 Sept 2026
- paperUnderstanding Emergent Abilities of Language Models from the Loss Perspective, NeurIPS · read 27 Sept 2026
- paperBeyond the Imitation Game: Quantifying and extrapolating the capabilities of language models, BIG-bench collaboration · read 27 Sept 2026
- paperTranscending Scaling Laws with 0.1% Extra Compute, Google Research · read 27 Sept 2026