Omni models
An omni model processes several input and output modalities within one multimodal model rather than routing everything through a text-only core.
An omni model is designed around several modalities at once. GPT-4o’s launch described one neural network trained end to end across text, vision and audio. Its stated input set included text, audio, images and video, while its outputs included text, audio and images.
That differs from a pipeline that transcribes audio, sends text through a language model and synthesizes speech afterward. OpenAI noted that such a text middle stage cannot directly observe tone, multiple speakers or background noise. End-to-end processing lets the model work from the original modalities, although each architecture still chooses how to encode and decode them.
Those choices vary. AnyGPT turns speech, text, images and music into discrete representations. Qwen2.5-Omni uses audio and visual encoders and interleaves audio with video timing. ImageBind instead demonstrates a joint embedding across six modalities.
The term describes a model design, not a deployment contract. At GPT-4o’s launch, only text and image inputs and text outputs were publicly released even though the research model covered more modalities. The launch also described the model as an early step with limitations still being explored.
A text-only middle stage can discard timing, tone and other information carried by the original signal.
Encode each incoming modality, fuse the resulting representations, then decode the requested output stream.
- 1 · encodeConvert each supported input modality into representations the shared model can process.
- 2 · alignPreserve the ordering or timing needed to relate information across modalities.
- 3 · reasonProcess the combined context inside the multimodal model.
- 4 · decodeGenerate through the text, speech or image output path supported by that model.
Architectures differ. AnyGPT uses discrete representations, while Qwen2.5-Omni uses audio and visual encoders.
| Who | What they ask | What it works with |
|---|---|---|
| Live assistant | “Discuss what the camera sees while listening to speech.” | Video frames, audio and the conversation |
| Media editor | “Follow a spoken instruction about an image.” | The instruction and image together |
| Accessibility tool | “Describe visual content as speech.” | Visual input and the requested audio response |
| Research benchmark | “Test whether one model transfers across modalities.” | Matched multimodal tasks and outputs |
- GPT-4o was trained end to end across text, vision and audio in one neural network.
- Qwen2.5-Omni accepts text, image, audio and video while generating streaming text and speech.
- AnyGPT represents speech, text, images and music as discrete sequences for unified processing.
- ImageBind learns one joint embedding across six modalities.
- A model's research capabilities do not guarantee that every deployed endpoint exposes every input and output modality.
- Combining modalities introduces modality-specific safety issues, including unauthorized voice generation.
- GPT-4o's launch described the model as an early step whose limitations were still being explored.
- OpenAI reported tasks where GPT-4 Turbo still outperformed GPT-4o.
Sources used
This explainer is written in original language. The links below support its factual claims.
- officialHello GPT-4o, OpenAI · read 28 Sept 2026
- paperQwen2.5-Omni Technical Report, Xu et al. · read 28 Sept 2026
- paperAnyGPT: Unified Multimodal LLM with Discrete Sequence Modeling, Zhan et al. · read 28 Sept 2026
- paperImageBind: One Embedding Space To Bind Them All, Girdhar et al. · read 28 Sept 2026
- officialGPT-4o System Card, OpenAI · read 28 Sept 2026