Concepts

Emergent abilities

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

Emergent abilities are task capabilities that appear absent in smaller models but present at a larger scale.

1 · What it is

The original definition focuses on the shape of measured task performance across scale. BIG-bench evaluated models spanning millions to hundreds of billions of parameters, and its authors saw three patterns on different tasks, steady gains, no gains, and sudden breakthroughs. The sudden jumps often came on tasks where the model must chain several parts together, or where the scoring rule is fragile.

But a jump in a chart is not automatically a jump in the underlying behavior. The metric critique showed that a nonlinear or discontinuous score can turn smooth output changes into an apparent cliff. With continuous metrics, the same fixed outputs can instead form a smooth and predictable curve.

The debate did not end there. Later loss-based work reported thresholds even under continuous metrics, and found that models with different parameter and data sizes can perform similarly at the same pre-training loss under controlled conditions. In the UL2R experiments, a changed training recipe made scores on a few BIG-bench tasks jump at 62B parameters, whereas the original PaLM only jumped at 540B. So emergence is still an open debate, not a settled law of scaling.

2 · Why it exists

A sudden benchmark jump can change what a model seems able to do and how teams forecast it.

Small models may look flatScores can stay near random across several scales and then rise sharply at a larger scale.
The metric can create the cliffAn all-or-nothing score such as exact match can make steady progress look like a sudden jump.
Scale is not the only axisChanging how a model is trained can make an apparent ability show up at a smaller model size.
3 · How it works

Test the same task across model scales, then inspect the measurement itself.

An apparent capability cliff can come from how model outputs are converted into a score.
  1. 1 · hold fixedPick one task and one model family, and save the outputs so every metric scores exactly the same answers.
  2. 2 · sweepEvaluate several model sizes rather than comparing only one small and one large model.
  3. 3 · inspectPlot both the task's headline metric and a continuous measure of partial progress when one is meaningful.
  4. 4 · retestCheck other tasks and other model families before treating one threshold as a general rule.

The key question is whether the behavior jumps or only the reported score does.

4 · Where it's used
WhoWhat they askWhat it works with
Evaluation team“Did reasoning truly appear at this scale?”Per-example outputs under exact and continuous metrics
Model researcher“Which checkpoint first exceeds chance?”Task score against model scale and pre-training loss
Safety team“Could a risky capability appear between releases?”Capability sweeps across intermediate checkpoints
Product lead“Will a larger model handle this workflow?”Repeated task evaluations under the intended prompt and metric
5 · What it solves, and what it doesn't
solves
  • The concept names sharp task-level changes that simple small-model extrapolation can miss.
  • Scale sweeps can reveal where performance first moves above random guessing.
  • Metric audits separate a true behavioral discontinuity from a thresholding artefact.
  • Loss-based analysis offers another axis for comparing models with different parameter and data sizes.
doesn't solve
  • A benchmark jump does not reveal the internal mechanism that produced it.
  • One sharp result does not prove that all models or prompts will show the same threshold.
  • A cliff seen under one metric may flatten into a smooth curve under another.
  • Parameter count alone does not predict task scores, since models of different sizes trained on different amounts of data can score alike.