Few-shot learning
Few-shot learning is the ability to perform a task from only a small number of labelled examples.
Few-shot learning names a situation, not a single method: a new task arrives with only a handful of labelled examples for each class. Those few examples are not enough to train a large model from nothing, because it would simply memorise them. So the system leans on earlier training across many other tasks, which left it with features, a starting point or a comparison rule it can reuse.
In language-model prompting, the support set appears as demonstrations in the input context and the model weights can remain fixed. Matching Networks instead learn to map a support set and query to a label without fine-tuning for the new class types. Prototypical Networks build one prototype per class and label a query by how close it sits to each prototype. MAML explicitly trains an initialization so a few gradient updates on new-task data can work well. Relation Networks learn a comparison function and can classify new classes without updating the network.
These mechanisms share a data constraint, not a single internal adaptation process. A credible evaluation therefore separates training classes or tasks from new ones and measures predictions on held-out queries. With prompts, just reordering the same examples can move accuracy anywhere between guessing and top benchmark scores, so one lucky set of examples proves little.
Many useful tasks have only a handful of examples because labels are expensive, rare or newly defined.
Pair a small support set with a new query, then use a pretrained adaptation rule to produce the answer.
- 1 · pretrainLearn reusable representations, an initialization or an in-context capability from many earlier examples or tasks.
- 2 · supportSupply a small labelled support set for the new task, next to an unlabelled query that needs an answer.
- 3 · adaptCondition on the examples, compare learned representations or take a small number of gradient steps.
- 4 · queryApply the resulting task-specific decision rule to an unseen input from the same task.
- 5 · evaluateTest with several different sets of examples, because in prompt-based few-shot learning the picked examples and their order can swing accuracy.
The defining constraint is few labelled examples; whether weights change depends on the method.
| Who | What they ask | What it works with |
|---|---|---|
| Language-model user | “Classify this request using these three examples.” | Demonstrations placed in the prompt |
| Wildlife researcher | “Recognise a newly monitored species from five images.” | A support set and pretrained visual representation |
| Medical team | “Adapt a detector for a rare finding.” | Performance across repeated small training samples |
| ML evaluator | “Does the method transfer to classes never used for training?” | Held-out tasks, classes and query examples |
- It reduces the amount of labelled task-specific data needed for adaptation.
- It can use knowledge learned across earlier tasks or broad pretraining.
- It lets someone describe a new task in plain text with a few worked examples, instead of retraining the model.
- It provides a common evaluation setting for prompt-based and meta-learned adaptation.
- It does not guarantee reliable generalization from a tiny or biased support set.
- It does not name one algorithm; different few-shot methods update different state or no weights at all.
- It does not make the choice of examples unimportant; with prompts, which examples you pick and their order can change accuracy.
- It does not mean zero training data; prior training supplies the structure that makes few-example adaptation possible.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperLanguage Models are Few-Shot Learners, NeurIPS · read 27 Sept 2026
- paperMatching Networks for One Shot Learning, NeurIPS · read 27 Sept 2026
- paperPrototypical Networks for Few-shot Learning, NeurIPS · read 27 Sept 2026
- paperModel-Agnostic Meta-Learning for Fast Adaptation of Deep Networks, Proceedings of Machine Learning Research · read 27 Sept 2026
- paperLearning to Compare: Relation Network for Few-Shot Learning, CVPR Open Access · read 27 Sept 2026
- paperCalibrate Before Use: Improving Few-shot Performance of Language Models, Proceedings of Machine Learning Research · read 27 Sept 2026