Concepts

Chat templates

3 min readintermediateUpdated 28 Sept 2026
1 · In one line

A chat template converts a structured list of messages into a formatted sequence.

1 · What it is

An application usually stores a conversation as objects with role and content keys. The template loops over the messages. Its output uses the format and control tokens expected by that model. The rendered result becomes one token sequence for an ordinary language model to continue.

Tool-capable templates can receive a list of functions through the tools argument. A multimodal chat template can be set with processor.chat_template.

The continue_final_message option continues the final message instead of starting a new one and removes its end-of-sequence tokens. Extra whitespace that was absent during model training can harm performance.

2 · Why it exists

Chat models can use different formats or control tokens.

Roles need encodingThe template must match the format used to train the model.
Formats differDifferent models may use different formats or control tokens.
Generation needs a boundarySome models require a final assistant header that signals where a new response should start.
3 · How it works

Render each structured message with the model's role markers and turn delimiters, append any generation marker, then tokenize once.

A chat template converts message objects into a formatted sequence.
  1. 1 · collectBuild an ordered list of messages with role and content fields.
  2. 2 · selectRead the Jinja template stored in the tokenizer's chat_template attribute.
  3. 3 · renderUse the model-specific format and control tokens.
  4. 4 · promptAdd tokens that indicate the start of an assistant response.
  5. 5 · tokenizeTokenize the rendered result without duplicating special tokens.

The correct template is the one that matches the model's training format.

4 · Where it's used
WhoWhat they askWhat it works with
Inference engineer“Why is this chat model continuing the user's message?”The generation prompt and assistant header
Model publisher“How should users format messages for this checkpoint?”The packaged chat_template.jinja file
Agent developer“How are tool schemas and tool results placed in the prompt?”Template branches and tool-message serialization
Multimodal developer“Where do image placeholders enter the conversation?”The processor template and media special tokens
5 · What it solves, and what it doesn't
solves
  • It accepts conversations structured as a list of dictionaries with role and content keys.
  • It handles model-specific formats and control tokens.
  • It can receive tools and retrieved documents for supported templates.
  • A template can be stored as a standalone Jinja file that is easy to inspect, edit and diff.
doesn't solve
  • Exact output tokens depend on the model.
  • It does not make every model emit the same tool-call format.
  • It does not prevent duplicated special tokens when rendered text is tokenized with special-token insertion enabled.
  • It does not parse every generated token back into structured messages unless a response parser is also defined.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. docsChat templates, Hugging Face Transformers · read 27 Sept 2026
  2. docsWriting a chat template, Hugging Face Transformers · read 27 Sept 2026
  3. docsExpanding Chat Templates with Tools and Documents, Hugging Face Transformers · read 28 Sept 2026
  4. docsMultimodal Chat Templates for Vision and Audio LLMs, Hugging Face Transformers · read 27 Sept 2026
  5. docsResponse Parsing, Hugging Face Transformers · read 27 Sept 2026