Concepts

Instruct models

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

An instruct model is a pretrained base version further fine-tuned on instructions and conversational data.

1 · What it is

An instruct model begins with a pretrained checkpoint and receives additional training on instructions or conversations.

FLAN instruction-tunes a pretrained model on multiple tasks phrased as instructions. InstructGPT adds a supervised stage and then reinforcement learning from ranked outputs.

Gemma instruction-tuned models are trained with a specific formatter that identifies roles in a conversation. Instruction tuning can help a model respond better to instructions. The InstructGPT paper reports that its models still make simple mistakes.

2 · Why it exists

Text completion and instruction following are different model uses.

Intent gapA base model suited to text completion is not ideal for tasks that require following instructions.
Task transferScaling instruction fine-tuning can improve generalization to unseen tasks.
Conversation formatGemma's formatter identifies roles in a conversation.
3 · How it works

Turn a base checkpoint into an instruction-following checkpoint.

Instruction tuning changes the checkpoint with examples of requests and desired responses.
  1. 1 · collectCollect demonstrations of desired model behavior.
  2. 2 · tuneFine-tune the pretrained model on those instruction examples.
  3. 3 · refineSome training pipelines add ranked outputs and human-feedback optimization.
  4. 4 · formatPresent user and model turns in the chat format expected by the checkpoint.
  5. 5 · respondUse the tuned model to respond better to instructions.

Instruct describes further tuning applied to a pretrained base version.

4 · Where it's used
WhoWhat they askWhat it works with
Product engineer“Which checkpoint should answer direct user requests?”Instruct variant
Model trainer“Which demonstrations should shape response behavior?”Instruction-tuning data
Inference engineer“Which role markers does this checkpoint expect?”Chat template
5 · What it solves, and what it doesn't
solves
  • It trains a pretrained model to respond better to natural-language instructions.
  • Instruction tuning can improve generalization to unseen tasks.
  • Chat formatting can identify conversational roles.
  • Chat formatting can delineate turns in a conversation.
doesn't solve
  • InstructGPT still makes simple mistakes.
  • Gemma instruction-tuned models are trained with a specific formatter.
  • Instruction tuning and preference optimization are separate stages in some pipelines.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsLLM prompting guide, Hugging Face · read 28 Sept 2026
  2. paperFinetuned Language Models Are Zero-Shot Learners, Wei et al. · read 28 Sept 2026
  3. paperTraining language models to follow instructions with human feedback, Ouyang et al. · read 28 Sept 2026
  4. paperScaling Instruction-Finetuned Language Models, Chung et al. · read 28 Sept 2026
  5. docsGemma formatting and system instructions, Google AI for Developers · read 28 Sept 2026