Concepts

Few-shot learning

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

Few-shot learning is the ability to perform a task from only a small number of labelled examples.

1 · What it is

Few-shot learning names a situation, not a single method: a new task arrives with only a handful of labelled examples for each class. Those few examples are not enough to train a large model from nothing, because it would simply memorise them. So the system leans on earlier training across many other tasks, which left it with features, a starting point or a comparison rule it can reuse.

In language-model prompting, the support set appears as demonstrations in the input context and the model weights can remain fixed. Matching Networks instead learn to map a support set and query to a label without fine-tuning for the new class types. Prototypical Networks build one prototype per class and label a query by how close it sits to each prototype. MAML explicitly trains an initialization so a few gradient updates on new-task data can work well. Relation Networks learn a comparison function and can classify new classes without updating the network.

These mechanisms share a data constraint, not a single internal adaptation process. A credible evaluation therefore separates training classes or tasks from new ones and measures predictions on held-out queries. With prompts, just reordering the same examples can move accuracy anywhere between guessing and top benchmark scores, so one lucky set of examples proves little.

2 · Why it exists

Many useful tasks have only a handful of examples because labels are expensive, rare or newly defined.

Little evidenceA conventional high-capacity model can overfit a tiny labelled set instead of learning what transfers.
New tasksThe examples may introduce classes or output rules that were absent from the original training task.
Several mechanismsPrompt demonstrations, metric comparison and rapid fine-tuning are different ways to use a few examples.
3 · How it works

Pair a small support set with a new query, then use a pretrained adaptation rule to produce the answer.

Few-shot learning transfers prior structure into a new task using only a small support set.
  1. 1 · pretrainLearn reusable representations, an initialization or an in-context capability from many earlier examples or tasks.
  2. 2 · supportSupply a small labelled support set for the new task, next to an unlabelled query that needs an answer.
  3. 3 · adaptCondition on the examples, compare learned representations or take a small number of gradient steps.
  4. 4 · queryApply the resulting task-specific decision rule to an unseen input from the same task.
  5. 5 · evaluateTest with several different sets of examples, because in prompt-based few-shot learning the picked examples and their order can swing accuracy.

The defining constraint is few labelled examples; whether weights change depends on the method.

4 · Where it's used
WhoWhat they askWhat it works with
Language-model user“Classify this request using these three examples.”Demonstrations placed in the prompt
Wildlife researcher“Recognise a newly monitored species from five images.”A support set and pretrained visual representation
Medical team“Adapt a detector for a rare finding.”Performance across repeated small training samples
ML evaluator“Does the method transfer to classes never used for training?”Held-out tasks, classes and query examples
5 · What it solves, and what it doesn't
solves
  • It reduces the amount of labelled task-specific data needed for adaptation.
  • It can use knowledge learned across earlier tasks or broad pretraining.
  • It lets someone describe a new task in plain text with a few worked examples, instead of retraining the model.
  • It provides a common evaluation setting for prompt-based and meta-learned adaptation.
doesn't solve
  • It does not guarantee reliable generalization from a tiny or biased support set.
  • It does not name one algorithm; different few-shot methods update different state or no weights at all.
  • It does not make the choice of examples unimportant; with prompts, which examples you pick and their order can change accuracy.
  • It does not mean zero training data; prior training supplies the structure that makes few-example adaptation possible.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperLanguage Models are Few-Shot Learners, NeurIPS · read 27 Sept 2026
  2. paperMatching Networks for One Shot Learning, NeurIPS · read 27 Sept 2026
  3. paperPrototypical Networks for Few-shot Learning, NeurIPS · read 27 Sept 2026
  4. paperModel-Agnostic Meta-Learning for Fast Adaptation of Deep Networks, Proceedings of Machine Learning Research · read 27 Sept 2026
  5. paperLearning to Compare: Relation Network for Few-Shot Learning, CVPR Open Access · read 27 Sept 2026
  6. paperCalibrate Before Use: Improving Few-shot Performance of Language Models, Proceedings of Machine Learning Research · read 27 Sept 2026