Concepts

Vision transformers

3 min readbeginnerUpdated 28 Sept 2026
1 · In one line

A vision transformer represents an image as a sequence of patches and processes that sequence with a transformer.

1 · What it is

ViT turns non-overlapping image patches into a token sequence. Position embeddings are added to patch embeddings. The patch sequence is sent through a standard Transformer encoder.

A learnable class token is placed at the front of the sequence. After the encoder, a classification head reads that token’s final state.

DeiT introduced a teacher-student strategy for transformers. Swin computes representations with shifted windows. A masked autoencoder can encode only the visible subset of patches.

2 · Why it exists

ViT reshapes an image into a sequence of flattened patches.

Sequence inputViT applies a transformer directly to a sequence of image patches.
Patch positionPosition embeddings are added to patch embeddings before the encoder.
Visual outputThe final class-token state serves as the image representation.
3 · How it works

Follow one image from pixels to a class score.

Patch projection converts the image grid into tokens. The transformer encoder processes those tokens as one sequence.
  1. 1 · splitThe image is divided into fixed-size patches.
  2. 2 · projectFlattened patches are mapped to patch embeddings.
  3. 3 · class tokenA learnable class embedding is prepended to the patch sequence.
  4. 4 · encodeThe patch sequence is sent through a standard Transformer encoder.
  5. 5 · classifyThe classification head reads the final class-token state.

The central conversion is image grid to token sequence.

4 · Where it's used
WhoWhat they askWhat it works with
Image classifier“Which category best matches this image?”Patch representations across the image
Detection system“Which regions contain objects?”Multi-scale visual representations
Self-supervised learner“What pixels belong in the masked patches?”Visible patch representations and mask tokens
5 · What it solves, and what it doesn't
solves
  • A pure transformer can operate directly on image-patch sequences.
  • Shifted windows can limit attention to local windows while connecting across windows.
  • Masked autoencoders can learn by reconstructing missing image patches.
doesn't solve
  • The original ViT results relied on pre-training with large amounts of data.
  • The original ViT paper evaluates image classification tasks.