Vision-language-action models
A vision-language-action model maps visual observations and a language instruction to robot actions.
A vision-language-action model, or VLA, extends multimodal understanding into control. It consumes visual observations and an instruction, then predicts actions a robot can execute. RT-2, for example, represents actions as text tokens.
RT-2 established a direct recipe: encode robot actions as text tokens and train those tokens alongside Internet-scale vision-language tasks. OpenVLA was trained on 970,000 real-world robot demonstrations.
Other designs use different action representations. Pi-zero builds a flow-matching architecture on a pretrained vision-language model. Octo can take language commands or goal images and can be fine-tuned for new sensory inputs and action spaces.
Vision-language pretraining is not enough by itself. RT-2 combines Internet-scale vision-language tasks with robot trajectory data. Octo’s results also show that new observation and action spaces can require fine-tuning.
A robot policy must connect visual and linguistic context to actions in the physical world.
Encode images, robot state and the instruction, then decode an action or action chunk.
- 1 · observeCollect visual observations and the robot's current state.
- 2 · instructSupply the task as a natural-language instruction or goal image.
- 3 · predictDecode the model representation into robot actions.
- 4 · executeUse the predicted controls for end-to-end robot control.
RT-2 writes actions as text tokens. Pi-zero builds a flow-matching architecture on a pretrained vision-language model.
| Who | What they ask | What it works with |
|---|---|---|
| Manipulation robot | “Put the named object in the bowl.” | Camera views, instruction and robot state |
| Mobile manipulator | “Carry laundry from the dryer to the table.” | Visual observations across the task |
| Robot researcher | “Adapt one policy to a new robot setup.” | Demonstrations from the new sensors and action space |
| Generalist policy | “Follow either a language command or a goal image.” | The chosen goal representation |
- RT-2 expresses robot actions as text tokens and trains them alongside vision-language tasks.
- OpenVLA was trained on 970,000 real-world robot demonstrations.
- Octo accepts language commands or goal images and can be fine-tuned to new sensor and action spaces.
- pi-zero builds a flow-matching architecture on a pretrained vision-language model.
- Vision-language pretraining does not remove the need for robot trajectory data.
- A policy trained on particular robots may need adaptation for a new sensor or action space.
- Robot learning still faces obstacles in data, generalization and robustness.
- Grounding language in real-world continuous sensor inputs remains a central robotics challenge.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperRT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control, Brohan et al. · read 28 Sept 2026
- paperOpenVLA: An Open-Source Vision-Language-Action Model, Kim et al. · read 28 Sept 2026
- paperOcto: An Open-Source Generalist Robot Policy, Octo Model Team et al. · read 28 Sept 2026
- paperpi-zero: A Vision-Language-Action Flow Model for General Robot Control, Physical Intelligence · read 28 Sept 2026
- paperPaLM-E: An Embodied Multimodal Language Model, Driess et al. · read 28 Sept 2026