Concepts

Vision-language models

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

A vision-language model can take visual data and text as input and produce text as output.

1 · What it is

A vision-language model links visual input with language. PaLI reports image captioning and visual question answering.

CLIP learns image and text representations for zero-shot transfer. BLIP-2 places a trainable Q-Former between a frozen image encoder and a frozen language model. LLaVA connects a vision encoder and an LLM. Flamingo accepts interleaved visual data and text.

Flamingo reports hallucinations and ungrounded guesses.

2 · Why it exists

BLIP-2 bridges a frozen image encoder and a frozen language model.

Visual featuresAn image encoder extracts information from pixels.
Language linkCLIP learns perception from supervision contained in natural language.
Open answersFlamingo accepts interleaved visual data and text and produces free-form text.
3 · How it works

Follow the trainable bridge in BLIP-2.

BLIP-2 connects a frozen image encoder to a frozen language model through a trainable Q-Former.
  1. 1 · seeA vision encoder turns the image into visual features.
  2. 2 · bridgeA learned component connects visual features to the language side.
  3. 3 · connectPlace the trainable Q-Former between the frozen image encoder and frozen language model.

Flamingo produces free-form text from interleaved visual data and text.

4 · Where it's used
WhoWhat they askWhat it works with
Accessibility team“What is shown in this image?”Visual features and caption generation
Document team“What does this chart say?”Image content and a written question
Search engineer“Which image best matches this phrase?”Image-text embedding similarity
5 · What it solves, and what it doesn't
solves
  • LLaVA connects a vision encoder and a language model.
  • PaLI reports image captioning and visual question-answering tasks.
  • CLIP connects image representations to language for flexible zero-shot transfer.
doesn't solve
  • Visual-language generation can still hallucinate or make ungrounded guesses.
  • Zero-shot visual performance can be weak on specialized or abstract tasks.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperLearning Transferable Visual Models From Natural Language Supervision, Radford et al., OpenAI · read 28 Sept 2026
  2. paperFlamingo: a Visual Language Model for Few-Shot Learning, Alayrac et al., DeepMind · read 28 Sept 2026
  3. paperBLIP-2, Li et al. · read 28 Sept 2026
  4. paperVisual Instruction Tuning, Liu et al. · read 28 Sept 2026
  5. paperPaLI: A Jointly-Scaled Multilingual Language-Image Model, Chen et al., Google Research · read 28 Sept 2026