Concepts

Supervised fine-tuning

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

Supervised fine-tuning keeps training an already-trained model on example questions paired with good answers, so it learns to answer that way.

1 · What it is

Many large language models first learn by guessing the next word on web pages. That gives them broad knowledge, but not the habit of doing what you ask. Supervised fine-tuning, or SFT, is a second round of training. You show the model example prompts, each paired with the answer you want, and nudge it towards giving answers like those.

Here is the mechanism. The model reads a prompt and the target answer, broken into small pieces called tokens. For each token of the answer, it scores how likely it thought that token was. A number called the loss measures how far off it was. Think of it as a score where lower is better. Training then adjusts the model’s weights, its internal settings, to shrink that loss. In Hugging Face’s SFT tool, only the answer is scored by default, not the prompt.

OpenAI’s InstructGPT used this recipe. OpenAI fine-tuned GPT-3 on answers written by human labellers, then improved it further with rankings of model outputs. Later, the LIMA study used only 1,000 carefully chosen examples. Its authors think SFT mostly teaches style and format. In their view, it brings out knowledge the model already had. Either way, keep some examples aside and test on them, to check that the tuned model really beats the base model.

2 · Why it exists

A model trained only to continue web text does not automatically do what people ask.

Wrong training goalGuessing the next word of a web page is a different goal from following a person's request helpfully and safely.
Unhelpful answersLarge language models can give answers that are untrue, harmful or simply not useful.
No proof it helpedWithout a fair test, you cannot tell whether the tuned model is better than the one you started with.
3 · How it works

Follow one example pair through a training step.

How supervised fine-tuning changes a language model A prompt and ideal response become tokens. A base model predicts response tokens. The prediction is compared with the ideal response, producing a loss that updates the model weights. A separate held-out prompt tests the tuned model, and its answer is compared with a target and the base model. Nothing from the test flows back into training. TRAINING · KNOWN-GOOD RESPONSE EVALUATION · ANSWER KEPT OUT OF TRAINING Training pair Base model Compare tokens Update Tuned model Held-out prompt Tuned model Score against target prompt + ideal response predicts the response tokens adjusts weights to cut the loss new weights not seen during weight updates generates a response compare with target and the base model target: approved guess: delayed cross-entropy loss USE THE UPDATED WEIGHTS NO TRAINING ARROW RETURNS
Training changes the model's weights using known-good answers. Testing uses examples the model never trained on.
  1. 1 · collectGather example prompts, each paired with the answer you want.
  2. 2 · formatPut each pair into the prompt-and-completion or chat format the training tool expects.
  3. 3 · compareScore how likely the model finds each token of the target answer, using a measure called cross-entropy loss.
  4. 4 · updateAdjust the model's weights to shrink the gap between its predictions and the target answers.
  5. 5 · testCheck the tuned model against the original on held-out examples that were kept out of training.

The known-good answer is the supervision; the weight update is the fine-tuning.

4 · Where it's used
WhoWhat they askWhat it works with
App builder“Reply in one fixed format”Example requests paired with correctly formatted replies
News app team“Label each article as business or entertainment”Articles paired with the correct label
Research team“Summarize chats in one fixed format”Chats paired with approved summaries
5 · What it solves, and what it doesn't
solves
  • It can make a model produce the style and content you want more reliably.
  • Official guides list uses such as classification, summarization, translation and fixed output formats.
  • A small set of carefully chosen, varied examples can go a long way.
  • It works best for a clearly defined task where labelled examples exist.
doesn't solve
  • It does not guarantee correct answers. InstructGPT, trained partly this way, could still ignore instructions and make up facts.
  • It does not help every task equally. FLAN gained less on tasks that were already plain sentence completion.
  • Adding more examples alone may not help; variety and answer quality mattered more in the LIMA study.
  • It does not prove itself. You still need to test the result against the base model.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperTraining language models to follow instructions with human feedback, Ouyang et al., OpenAI · read 28 Sept 2026
  2. paperFinetuned Language Models Are Zero-Shot Learners, Wei et al., Google Research · read 28 Sept 2026
  3. paperLIMA: Less Is More for Alignment, Zhou et al., Meta AI · read 28 Sept 2026
  4. docsSFT Trainer, Hugging Face · read 28 Sept 2026
  5. docsSupervised fine-tuning, OpenAI · read 28 Sept 2026
  6. docsAbout supervised fine-tuning for Gemini models, Google Cloud · read 28 Sept 2026