Self-supervised learning
Self-supervised learning trains a model on raw data by hiding part of each example and asking the model to predict it, so the data supplies its own labels.
Self-supervised learning is a way to train a model without anyone writing labels. You take ordinary data, hide part of it, and ask the model to predict the missing piece from what is left. The hidden piece is the answer, so the data labels itself. In effect, a problem with no labels is turned into one that can be marked right or wrong.
The trick comes in several forms. BERT, a Google language model, learned by filling in masked words in plain text from Wikipedia. OpenAI’s GPT-2 learned by predicting the next word in 40GB of internet text. Pictures work too: MAE hides 75% of an image’s patches and rebuilds the missing pixels, while SimCLR shows a model two edited copies of the same photo and trains it to recognise them as a pair.
This is usually the first of two stages. In pretraining, the model picks up general patterns from a huge unlabelled pile. In fine-tuning, a far smaller labelled dataset teaches it one job, such as judging whether a review is positive. That places it between two other families. Like supervised learning, it scores every guess against an answer. Like unsupervised learning, it needs no human labels, which is why Meta’s researchers call the old name “unsupervised” misleading here.
Supervised learning needs answers written by people, and there are never enough.
Follow one sentence through a training step.
- 1 · collectGather a large amount of raw data, such as web text or photos, with no labels added by people.
- 2 · hideHide part of each example, such as some of the words in a sentence, and keep the hidden part aside as the answer.
- 3 · predictThe model fills in the gap using only the parts it can still see.
- 4 · correctThe guess is checked against the hidden original, and the model is adjusted so it gets closer next time.
- 5 · adaptAfterwards, the pretrained model is fine-tuned on a small labelled dataset for one specific task.
The hidden part is the label. The data writes its own answer key.
| Who | What they ask | What it works with |
|---|---|---|
| Chatbot maker | “Can the model continue any piece of text sensibly?” | Huge amounts of web text, learning to predict each next word |
| Search team | “What does this query mean, given every word around it?” | A text model pretrained on masked words, then fine-tuned |
| Social network safety team | “Is this post hate speech, even in a language with few labelled examples?” | A multilingual text model pretrained without labels |
| Photo app team | “Can we sort pictures well using only a few labelled ones?” | Unlabelled images, learned by hiding patches or matching edited copies |
- Lets models learn from far more data than people could ever label.
- Cuts the number of labels a task needs; SimCLR reached 85.8% top-5 accuracy on ImageNet using 1% of the labels.
- One pretrained model can be fine-tuned for many different tasks with little change to its design.
- Pretrained models usually do considerably better than models trained only on labelled examples.
- It rarely finishes the job alone; a model is usually still fine-tuned on labelled data for a specific task.
- Pretraining is expensive; OpenAI's 2018 GPT needed a month on 8 GPUs for this step.
- The model learns what is in its data, including gaps, errors and biases.
- Filling in blanks is harder for images and video, where a missing piece could look countless different ways.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialSelf-supervised learning: The dark matter of intelligence, Meta AI · read 27 Sept 2026
- officialAdvancing Self-Supervised and Semi-Supervised Learning with SimCLR, Google Research · read 27 Sept 2026
- officialOpen Sourcing BERT: State-of-the-Art Pre-training for Natural Language Processing, Google Research · read 27 Sept 2026
- paperBERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, arXiv (Google AI Language) · read 27 Sept 2026
- officialBetter language models and their implications, OpenAI · read 27 Sept 2026
- officialImproving language understanding with unsupervised learning, OpenAI · read 27 Sept 2026
- paperMasked Autoencoders Are Scalable Vision Learners, arXiv (Meta AI Research, FAIR) · read 27 Sept 2026
- paperA Simple Framework for Contrastive Learning of Visual Representations, arXiv (Google Research) · read 27 Sept 2026
- docsMachine Learning Glossary, Google for Developers · read 27 Sept 2026