Robotics Transformer 2 (RT-2)
Introduction
Robotic Transformer 2 (RT-2) is a novel vision-language-action (VLA) model that learns from both web and robotics data, and translates this knowledge into generalised instructions for robotic control
High capacity VLMs (Vision Language Model) are trained on web-scale datasets. This makes them good at recognising visual or language patterns across different languages.
For robots, achieving similar level of competency would require robot data first hand, across:
- every object
- every environment
- every task
- every situation
RT-2 is novel VLA model that learns from web and robotics data and translates this knowledge into generalised instructions for robotic control while maintaining web-scale capabilities.
RT-1 trained on multi-task demonstrations, which can learn combinations of tasks and objects seen in the robotic data.
- Builds on this.
- Uses the robot demonstration data collected with 13 robots over 17 months in an office kitchen environment.
RT-2 shows:
- improved generalisation capabilities,
- semantic and visual understanding beyond exposed robotic data,
- ability to interpret new commands and respond to user commands by performing rudimentary reasoning (reasoning about object categories or high-level descriptions).
Incorporates chain-of-thought reasoning, allowing RT-2 to perform multi-stage semantic reasoning.
VLMs -> Robotics
RT-2 builds on VLMs that take >1 images as input and produce sequence of tokens that represent natural language text.
RT-2 uses Pathways Language and Image model (PaLI-X) and Pathways Language model Embodied (PaLM-E) as its backbone.
Robotic actions => Tokens; Actions are described as strings that can be processed by standard natural language tokenisers.

- This string could be represented with token numbers eg: "1 128 91 241 5 101 127 217"
RT-2 Architecture and Training
Co-fine-tune a pre-trained VLM on robotics and web data.
Resulting model takes in robot camera images and directly predicts actions for robot to perform.
Generalisation and Emergent Skills
6000 robotic trials - qualitative and quantitative experiments
Three categories of skills:
- Symbol understanding,
- Reasoning,
- Human Recognition.
When asked to perform tasks not seen in robotic data, requires translated knowledge from web-data to operate.
- "put strawberry into the correct bowl"
- "pick land animal"
- "pick animal with different colour"
- etc.

- Observed more than 3x improvement in generalisation performance compared to RT-1 and Visual Cortex (VC-1)
- 62% improved performance on previously unseen scenarios
Probed RT-2 to combine robotic control with chain-of-thought reasoning to enable long-horizon planning.
- Variant of RT-2 tuned to few hundred gradient steps to increase ability to use language and actions jointly.
- Augmented data to include additional "Plan" step: (1) Describe purpose of action; (2) "Action" in action tokens.

RT-2 can use both image and text commands.
Summary
VLMs can be transformed into powerful VLA models to directly control robot by combining VLM pre-training with robotic data.
PaLM-E