Instruct models
An instruct model is a pretrained base version further fine-tuned on instructions and conversational data.
An instruct model begins with a pretrained checkpoint and receives additional training on instructions or conversations.
FLAN instruction-tunes a pretrained model on multiple tasks phrased as instructions. InstructGPT adds a supervised stage and then reinforcement learning from ranked outputs.
Gemma instruction-tuned models are trained with a specific formatter that identifies roles in a conversation. Instruction tuning can help a model respond better to instructions. The InstructGPT paper reports that its models still make simple mistakes.
Text completion and instruction following are different model uses.
Turn a base checkpoint into an instruction-following checkpoint.
- 1 · collectCollect demonstrations of desired model behavior.
- 2 · tuneFine-tune the pretrained model on those instruction examples.
- 3 · refineSome training pipelines add ranked outputs and human-feedback optimization.
- 4 · formatPresent user and model turns in the chat format expected by the checkpoint.
- 5 · respondUse the tuned model to respond better to instructions.
Instruct describes further tuning applied to a pretrained base version.
| Who | What they ask | What it works with |
|---|---|---|
| Product engineer | “Which checkpoint should answer direct user requests?” | Instruct variant |
| Model trainer | “Which demonstrations should shape response behavior?” | Instruction-tuning data |
| Inference engineer | “Which role markers does this checkpoint expect?” | Chat template |
- It trains a pretrained model to respond better to natural-language instructions.
- Instruction tuning can improve generalization to unseen tasks.
- Chat formatting can identify conversational roles.
- Chat formatting can delineate turns in a conversation.
- InstructGPT still makes simple mistakes.
- Gemma instruction-tuned models are trained with a specific formatter.
- Instruction tuning and preference optimization are separate stages in some pipelines.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsLLM prompting guide, Hugging Face · read 28 Sept 2026
- paperFinetuned Language Models Are Zero-Shot Learners, Wei et al. · read 28 Sept 2026
- paperTraining language models to follow instructions with human feedback, Ouyang et al. · read 28 Sept 2026
- paperScaling Instruction-Finetuned Language Models, Chung et al. · read 28 Sept 2026
- docsGemma formatting and system instructions, Google AI for Developers · read 28 Sept 2026