Chat templates
A chat template converts a structured list of messages into a formatted sequence.
An application usually stores a conversation as objects with role and content keys. The template loops over the messages. Its output uses the format and control tokens expected by that model. The rendered result becomes one token sequence for an ordinary language model to continue.
Tool-capable templates can receive a list of functions through the tools argument. A multimodal chat template can be set with processor.chat_template.
The continue_final_message option continues the final message instead of starting a new one and removes its end-of-sequence tokens. Extra whitespace that was absent during model training can harm performance.
Chat models can use different formats or control tokens.
Render each structured message with the model's role markers and turn delimiters, append any generation marker, then tokenize once.
- 1 · collectBuild an ordered list of messages with role and content fields.
- 2 · selectRead the Jinja template stored in the tokenizer's chat_template attribute.
- 3 · renderUse the model-specific format and control tokens.
- 4 · promptAdd tokens that indicate the start of an assistant response.
- 5 · tokenizeTokenize the rendered result without duplicating special tokens.
The correct template is the one that matches the model's training format.
| Who | What they ask | What it works with |
|---|---|---|
| Inference engineer | “Why is this chat model continuing the user's message?” | The generation prompt and assistant header |
| Model publisher | “How should users format messages for this checkpoint?” | The packaged chat_template.jinja file |
| Agent developer | “How are tool schemas and tool results placed in the prompt?” | Template branches and tool-message serialization |
| Multimodal developer | “Where do image placeholders enter the conversation?” | The processor template and media special tokens |
- It accepts conversations structured as a list of dictionaries with role and content keys.
- It handles model-specific formats and control tokens.
- It can receive tools and retrieved documents for supported templates.
- A template can be stored as a standalone Jinja file that is easy to inspect, edit and diff.
- Exact output tokens depend on the model.
- It does not make every model emit the same tool-call format.
- It does not prevent duplicated special tokens when rendered text is tokenized with special-token insertion enabled.
- It does not parse every generated token back into structured messages unless a response parser is also defined.
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsChat templates, Hugging Face Transformers · read 27 Sept 2026
- docsWriting a chat template, Hugging Face Transformers · read 27 Sept 2026
- docsExpanding Chat Templates with Tools and Documents, Hugging Face Transformers · read 28 Sept 2026
- docsMultimodal Chat Templates for Vision and Audio LLMs, Hugging Face Transformers · read 27 Sept 2026
- docsResponse Parsing, Hugging Face Transformers · read 27 Sept 2026