Vision transformers
A vision transformer represents an image as a sequence of patches and processes that sequence with a transformer.
ViT turns non-overlapping image patches into a token sequence. Position embeddings are added to patch embeddings. The patch sequence is sent through a standard Transformer encoder.
A learnable class token is placed at the front of the sequence. After the encoder, a classification head reads that token’s final state.
DeiT introduced a teacher-student strategy for transformers. Swin computes representations with shifted windows. A masked autoencoder can encode only the visible subset of patches.
ViT reshapes an image into a sequence of flattened patches.
Follow one image from pixels to a class score.
- 1 · splitThe image is divided into fixed-size patches.
- 2 · projectFlattened patches are mapped to patch embeddings.
- 3 · class tokenA learnable class embedding is prepended to the patch sequence.
- 4 · encodeThe patch sequence is sent through a standard Transformer encoder.
- 5 · classifyThe classification head reads the final class-token state.
The central conversion is image grid to token sequence.
| Who | What they ask | What it works with |
|---|---|---|
| Image classifier | “Which category best matches this image?” | Patch representations across the image |
| Detection system | “Which regions contain objects?” | Multi-scale visual representations |
| Self-supervised learner | “What pixels belong in the masked patches?” | Visible patch representations and mask tokens |
- A pure transformer can operate directly on image-patch sequences.
- Shifted windows can limit attention to local windows while connecting across windows.
- Masked autoencoders can learn by reconstructing missing image patches.
- The original ViT results relied on pre-training with large amounts of data.
- The original ViT paper evaluates image classification tasks.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperAn Image is Worth 16x16 Words Transformers for Image Recognition at Scale, Dosovitskiy et al. · read 28 Sept 2026
- paperTraining data-efficient image transformers and distillation through attention, Touvron et al. · read 28 Sept 2026
- paperSwin Transformer Hierarchical Vision Transformer using Shifted Windows, Liu et al. · read 28 Sept 2026
- paperMasked Autoencoders Are Scalable Vision Learners, He et al. · read 28 Sept 2026
- docsVision Transformer ViT, Hugging Face · read 28 Sept 2026