Supervised fine-tuning
Supervised fine-tuning keeps training an already-trained model on example questions paired with good answers, so it learns to answer that way.
Many large language models first learn by guessing the next word on web pages. That gives them broad knowledge, but not the habit of doing what you ask. Supervised fine-tuning, or SFT, is a second round of training. You show the model example prompts, each paired with the answer you want, and nudge it towards giving answers like those.
Here is the mechanism. The model reads a prompt and the target answer, broken into small pieces called tokens. For each token of the answer, it scores how likely it thought that token was. A number called the loss measures how far off it was. Think of it as a score where lower is better. Training then adjusts the model’s weights, its internal settings, to shrink that loss. In Hugging Face’s SFT tool, only the answer is scored by default, not the prompt.
OpenAI’s InstructGPT used this recipe. OpenAI fine-tuned GPT-3 on answers written by human labellers, then improved it further with rankings of model outputs. Later, the LIMA study used only 1,000 carefully chosen examples. Its authors think SFT mostly teaches style and format. In their view, it brings out knowledge the model already had. Either way, keep some examples aside and test on them, to check that the tuned model really beats the base model.
A model trained only to continue web text does not automatically do what people ask.
Follow one example pair through a training step.
- 1 · collectGather example prompts, each paired with the answer you want.
- 2 · formatPut each pair into the prompt-and-completion or chat format the training tool expects.
- 3 · compareScore how likely the model finds each token of the target answer, using a measure called cross-entropy loss.
- 4 · updateAdjust the model's weights to shrink the gap between its predictions and the target answers.
- 5 · testCheck the tuned model against the original on held-out examples that were kept out of training.
The known-good answer is the supervision; the weight update is the fine-tuning.
| Who | What they ask | What it works with |
|---|---|---|
| App builder | “Reply in one fixed format” | Example requests paired with correctly formatted replies |
| News app team | “Label each article as business or entertainment” | Articles paired with the correct label |
| Research team | “Summarize chats in one fixed format” | Chats paired with approved summaries |
- It can make a model produce the style and content you want more reliably.
- Official guides list uses such as classification, summarization, translation and fixed output formats.
- A small set of carefully chosen, varied examples can go a long way.
- It works best for a clearly defined task where labelled examples exist.
- It does not guarantee correct answers. InstructGPT, trained partly this way, could still ignore instructions and make up facts.
- It does not help every task equally. FLAN gained less on tasks that were already plain sentence completion.
- Adding more examples alone may not help; variety and answer quality mattered more in the LIMA study.
- It does not prove itself. You still need to test the result against the base model.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperTraining language models to follow instructions with human feedback, Ouyang et al., OpenAI · read 28 Sept 2026
- paperFinetuned Language Models Are Zero-Shot Learners, Wei et al., Google Research · read 28 Sept 2026
- paperLIMA: Less Is More for Alignment, Zhou et al., Meta AI · read 28 Sept 2026
- docsSFT Trainer, Hugging Face · read 28 Sept 2026
- docsSupervised fine-tuning, OpenAI · read 28 Sept 2026
- docsAbout supervised fine-tuning for Gemini models, Google Cloud · read 28 Sept 2026