Robotics foundation models
A robotics foundation model is pretrained across many tasks, environments or robot bodies so it can serve as a starting point for new robot policies.
Robotics foundation models aim to reuse experience across physical systems. Open X-Embodiment begins from the problem that conventional robot learning often trains a separate model for every application, robot and environment. Its alternative is a standardized dataset and an X-robot policy that can be adapted to new settings.
The training mixture is heterogeneous. Demonstrations may come from different robot bodies, sensors, action spaces, tasks and environments. Open X-Embodiment assembled data from 22 robots. RoboCat consumed action-labelled visual experience from robot arms with varying observations and actions. Gato used one network and one set of weights across language, games and real-robot control.
The models do not share one architecture. PaLM-E incorporates continuous sensor inputs into a language model. RT-X learns from standardized cross-robot trajectories. Pi-zero builds a flow-matching architecture on a pretrained vision-language model.
The expected payoff is adaptation. RoboCat reported adapting to target tasks and robots with 100 to 1,000 examples. Open X-Embodiment reported positive transfer across robot platforms. Octo’s authors still identify diverse sensors and action spaces as requirements for broadly applicable policies.
Training a separate policy for every robot, task and environment prevents experience from being reused.
Standardize diverse robot experience, pretrain one broad policy, then adapt it to a target robot and task.
- 1 · collectGather action-labelled experience across robots, tasks and environments.
- 2 · standardizeExpress heterogeneous observations and actions in a training format the model can consume.
- 3 · pretrainTrain one policy on the combined multi-task and multi-embodiment mixture.
- 4 · adaptFine-tune the policy for a target setup with new sensory inputs or actions.
Gato uses one network and one set of weights. Pi-zero builds flow matching on a pretrained vision-language model.
| Who | What they ask | What it works with |
|---|---|---|
| Robotics lab | “Reuse demonstrations collected on another arm.” | Standardized trajectories from several embodiments |
| Manipulation team | “Add a new pick-and-place task with fewer new demonstrations.” | The base policy plus target-task examples |
| Mobile robot | “Combine language, vision and robot state for planning.” | Multimodal embodied observations |
| Dataset consortium | “Train an X-robot policy across institutions.” | A shared cross-embodiment dataset |
- Open X-Embodiment assembled data from 22 different robots.
- PaLM-E reported positive transfer from joint training across language, vision and visual-language domains.
- RoboCat adapts to new tasks and robots with 100 to 1,000 target examples in its reported experiments.
- Gato uses one network and one set of weights across text, games and real-robot control.
- A generalist policy still has to handle diverse sensors and action spaces.
- Robot learning still faces obstacles in data, generalization and robustness.
- A target setup with new sensory inputs or actions may still require fine-tuning.
- RoboCat's reported evaluation covered three real robot embodiments, not every robot body.
Sources used
This explainer is written in original language. The links below support its factual claims.
- paperOpen X-Embodiment: Robotic Learning Datasets and RT-X Models, Open X-Embodiment Collaboration · read 28 Sept 2026
- paperPaLM-E: An Embodied Multimodal Language Model, Driess et al. · read 28 Sept 2026
- paperRoboCat: A Self-Improving Generalist Agent for Robotic Manipulation, Bousmalis et al. · read 28 Sept 2026
- paperA Generalist Agent, Reed et al. · read 28 Sept 2026
- paperpi-zero: A Vision-Language-Action Flow Model for General Robot Control, Physical Intelligence · read 28 Sept 2026
- paperOcto: An Open-Source Generalist Robot Policy, Octo Model Team et al. · read 28 Sept 2026