Concepts
Multimodal models
1 · In one line
A multimodal model can process more than one type of input, such as text, images, audio or video.
1 · What it is
A multimodal model handles more than one kind of input. Google lists text, audio and video as examples of modalities.
Gemini was trained jointly across text, image, audio and video. ImageBind aligned six modalities in one embedding space. Flamingo processes interleaved visual data and text before generating text.
GPT-4’s report describes image and text inputs with text output.
Documents can combine text, photographs, diagrams or screenshots.
Mixed evidenceA document may contain text, photographs, diagrams or screenshots.
Linked meaningImageBind aligned six modalities in one embedding space.
Combined promptGPT-4 accepts prompts containing both images and text.
Follow several input types into one supported model.
- 1 · collectSupply one or more input types supported by the model.
- 2 · encodeIn Flamingo, a vision encoder and Perceiver Resampler extract visual tokens.
- 3 · combineGemini is trained jointly across text, image, audio and video.
- 4 · respondGPT-4 accepts image and text inputs and produces text output.
Multimodal means that more than one input type can appear in a prompt.
| Who | What they ask | What it works with |
|---|---|---|
| Document analyst | “What does this chart say alongside its caption?” | Image and text inputs |
| Media team | “What happens in this video clip?” | Video and audio inputs |
| Accessibility team | “Can this image be described in words?” | Visual input and text output support |
solves
- A single request can include more than one supported modality.
- Gemini is jointly trained across text, image, audio and video.
- Shared embeddings can align several modalities in one space.
doesn't solve
- Google documents that adding many images to a request increases response latency.
6 · Go deeper
Sources used
This explainer is written in original language. The links below support its factual claims.
- docsGenerate content with the Gemini API, Google Cloud · read 28 Sept 2026
- paperGPT-4 Technical Report, OpenAI · read 28 Sept 2026
- paperGemini: A Family of Highly Capable Multimodal Models, Gemini Team, Google · read 28 Sept 2026
- paperImageBind: One Embedding Space To Bind Them All, Girdhar et al., Meta AI · read 28 Sept 2026
- paperFlamingo: a Visual Language Model for Few-Shot Learning, Alayrac et al., DeepMind · read 28 Sept 2026