Concepts

Semi-supervised learning

5 min readintermediateUpdated 28 Sept 2026
1 · In one line

Semi-supervised learning trains a model on a few labelled examples plus many unlabelled ones, so the unlabelled data can help when answers are scarce.

1 · What it is

Semi-supervised learning trains a model on two kinds of data at once: a small set of examples that come with the right answer, called a label, and a much larger set that come without one. It exists because labels are the costly part. Someone has to do the labelling, which costs human work and wages, and in medicine a single label can require an invasive test.

One method is self-training. A model trained on the labelled examples guesses labels for the unlabelled ones, and only its most confident guesses, called pseudo-labels, join the training set before it trains again. A second family, consistency training, teaches the model to give the same answer for an example and an altered copy, such as a rotated photo or a reworded sentence. On a sentiment task built from IMDb data, a Google model of this kind trained with 20 labelled examples plus 50,000 unlabelled ones beat an earlier model trained on 25,000 labelled examples. FixMatch joins both ideas: a confident guess on a lightly altered image becomes the target for a heavily altered version of the same image.

All of this rests on an assumption. The unlabelled examples must show real structure, such as tight groups of similar examples that belong to the same class. When that breaks, the method can backfire. A model can end up learning its own wrong guesses, a trap called confirmation bias, and unlabelled data from classes the task never covered can leave a model worse than no unlabelled data at all. In its entry on self-supervised learning, Google’s glossary calls that method a semi-supervised approach too, one that makes stand-in labels out of the unlabelled examples themselves.

2 · Why it exists

Answers are the expensive part of training data.

Labels cost peopleSomeone has to sit and label every example, labellers must be paid, and in medicine one label can mean an invasive test.
Raw data is plentifulThe approach pays off when answers are expensive to get but examples without answers are plentiful.
A few labels misleadGive a model a handful of labelled points and it has little to go on when deciding where one class ends and the next begins.
3 · How it works

Follow one round of self-training.

Numbers from a scikit-learn test on a breast-cancer dataset with 50 of 569 examples labelled, where a threshold near 0.7 worked best.
  1. 1 · trainTrain an ordinary supervised model on the small labelled set.
  2. 2 · guessUse that model to predict a label, with a confidence score, for every unlabelled example.
  3. 3 · filterKeep only the guesses whose confidence clears a threshold; these become pseudo-labels.
  4. 4 · retrainAdd the pseudo-labelled examples to the training set and train again.
  5. 5 · repeatLoop until no new guesses pass the threshold or the model stops improving.

The confidence threshold carries the method: set it too low and wrong labels get in, too high and nothing new is learned.

4 · Where it's used
WhoWhat they askWhat it works with
Hospital imaging team“Can we train a scan classifier when specialists have read only a few scans?”A small set of read scans and a large archive of unread ones
Review site“Is this review positive or negative?”A few rated reviews and many unrated ones
Photo app“Which of these pictures show a dog?”A few tagged photos and a flood of untagged uploads
Support desk“Which team should handle this ticket?”A few routed tickets and a backlog of unrouted ones
5 · What it solves, and what it doesn't
solves
  • Puts unlabelled data to work that plain supervised learning would ignore.
  • Can reach high accuracy with very few labels, such as FixMatch's 88.61% on an image benchmark with 4 labels per class.
  • Wraps around existing tools; scikit-learn's self-training works with any classifier that outputs probabilities.
  • Can help even when labels are plentiful, as Noisy Student showed on ImageNet with 300 million unlabelled images.
doesn't solve
  • It still needs some labelled examples; it cuts labelling, it does not remove it.
  • A model can learn its own wrong pseudo-labels as if they were true, a trap called confirmation bias.
  • Unlabelled data from classes outside the task can leave a model worse than using none at all.
  • Reported gains shrink when the labels-only baseline gets the same tuning effort.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026
  2. docs1.14. Semi-supervised learning (User Guide), scikit-learn · read 27 Sept 2026
  3. docsEffect of varying threshold for self-training, scikit-learn · read 27 Sept 2026
  4. paperFixMatch: Simplifying Semi-Supervised Learning with Consistency and Confidence, NeurIPS 2020 (Google Research, via arXiv) · read 27 Sept 2026
  5. paperRealistic Evaluation of Deep Semi-Supervised Learning Algorithms, NeurIPS 2018 (Google Brain, via arXiv) · read 27 Sept 2026
  6. paperPseudo-Labeling and Confirmation Bias in Deep Semi-Supervised Learning, arXiv (Insight Centre for Data Analytics, Dublin City University) · read 27 Sept 2026
  7. paperSelf-training with Noisy Student improves ImageNet classification, CVPR 2020 (Google Research, via arXiv) · read 27 Sept 2026
  8. officialAdvancing Semi-supervised Learning with Unsupervised Data Augmentation, Google Research blog · read 27 Sept 2026