Concepts

Robotics foundation models

4 min readintermediateUpdated 28 Sept 2026
1 · In one line

A robotics foundation model is pretrained across many tasks, environments or robot bodies so it can serve as a starting point for new robot policies.

1 · What it is

Robotics foundation models aim to reuse experience across physical systems. Open X-Embodiment begins from the problem that conventional robot learning often trains a separate model for every application, robot and environment. Its alternative is a standardized dataset and an X-robot policy that can be adapted to new settings.

The training mixture is heterogeneous. Demonstrations may come from different robot bodies, sensors, action spaces, tasks and environments. Open X-Embodiment assembled data from 22 robots. RoboCat consumed action-labelled visual experience from robot arms with varying observations and actions. Gato used one network and one set of weights across language, games and real-robot control.

The models do not share one architecture. PaLM-E incorporates continuous sensor inputs into a language model. RT-X learns from standardized cross-robot trajectories. Pi-zero builds a flow-matching architecture on a pretrained vision-language model.

The expected payoff is adaptation. RoboCat reported adapting to target tasks and robots with 100 to 1,000 examples. Open X-Embodiment reported positive transfer across robot platforms. Octo’s authors still identify diverse sensors and action spaces as requirements for broadly applicable policies.

2 · Why it exists

Training a separate policy for every robot, task and environment prevents experience from being reused.

Fragmented dataRobot demonstrations come from different sensors, action spaces and embodiments.
Data obstacleGeneralist robot policies still face major obstacles in data, generalization and robustness.
New setupA reusable base model can be adapted rather than training a new policy from scratch.
3 · How it works

Standardize diverse robot experience, pretrain one broad policy, then adapt it to a target robot and task.

The foundation comes from reusable training across tasks and embodiments, followed by targeted adaptation.
  1. 1 · collectGather action-labelled experience across robots, tasks and environments.
  2. 2 · standardizeExpress heterogeneous observations and actions in a training format the model can consume.
  3. 3 · pretrainTrain one policy on the combined multi-task and multi-embodiment mixture.
  4. 4 · adaptFine-tune the policy for a target setup with new sensory inputs or actions.

Gato uses one network and one set of weights. Pi-zero builds flow matching on a pretrained vision-language model.

4 · Where it's used
WhoWhat they askWhat it works with
Robotics lab“Reuse demonstrations collected on another arm.”Standardized trajectories from several embodiments
Manipulation team“Add a new pick-and-place task with fewer new demonstrations.”The base policy plus target-task examples
Mobile robot“Combine language, vision and robot state for planning.”Multimodal embodied observations
Dataset consortium“Train an X-robot policy across institutions.”A shared cross-embodiment dataset
5 · What it solves, and what it doesn't
solves
  • Open X-Embodiment assembled data from 22 different robots.
  • PaLM-E reported positive transfer from joint training across language, vision and visual-language domains.
  • RoboCat adapts to new tasks and robots with 100 to 1,000 target examples in its reported experiments.
  • Gato uses one network and one set of weights across text, games and real-robot control.
doesn't solve
  • A generalist policy still has to handle diverse sensors and action spaces.
  • Robot learning still faces obstacles in data, generalization and robustness.
  • A target setup with new sensory inputs or actions may still require fine-tuning.
  • RoboCat's reported evaluation covered three real robot embodiments, not every robot body.
6 · Go deeper

Sources used

This explainer is written in original language. The links below support its factual claims.

  1. paperOpen X-Embodiment: Robotic Learning Datasets and RT-X Models, Open X-Embodiment Collaboration · read 28 Sept 2026
  2. paperPaLM-E: An Embodied Multimodal Language Model, Driess et al. · read 28 Sept 2026
  3. paperRoboCat: A Self-Improving Generalist Agent for Robotic Manipulation, Bousmalis et al. · read 28 Sept 2026
  4. paperA Generalist Agent, Reed et al. · read 28 Sept 2026
  5. paperpi-zero: A Vision-Language-Action Flow Model for General Robot Control, Physical Intelligence · read 28 Sept 2026
  6. paperOcto: An Open-Source Generalist Robot Policy, Octo Model Team et al. · read 28 Sept 2026