Robotics Transformer 2 (RT-2)

Introduction

Robotic Transformer 2 (RT-2) is a novel vision-language-action (VLA) model that learns from both web and robotics data, and translates this knowledge into generalised instructions for robotic control

High capacity VLMs (Vision Language Model) are trained on web-scale datasets. This makes them good at recognising visual or language patterns across different languages.

For robots, achieving similar level of competency would require robot data first hand, across:

RT-2 is novel VLA model that learns from web and robotics data and translates this knowledge into generalised instructions for robotic control while maintaining web-scale capabilities.

RT-1 trained on multi-task demonstrations, which can learn combinations of tasks and objects seen in the robotic data.

RT-2 shows:

Incorporates chain-of-thought reasoning, allowing RT-2 to perform multi-stage semantic reasoning.

VLMs -> Robotics

RT-2 builds on VLMs that take >1 images as input and produce sequence of tokens that represent natural language text.

RT-2 uses Pathways Language and Image model (PaLI-X) and Pathways Language model Embodied (PaLM-E) as its backbone.

Robotic actions => Tokens; Actions are described as strings that can be processed by standard natural language tokenisers.

rt21.png

RT-2 Architecture and Training

Co-fine-tune a pre-trained VLM on robotics and web data.

Resulting model takes in robot camera images and directly predicts actions for robot to perform.

Generalisation and Emergent Skills

6000 robotic trials - qualitative and quantitative experiments

Three categories of skills:

When asked to perform tasks not seen in robotic data, requires translated knowledge from web-data to operate.

rt22.webp

Probed RT-2 to combine robotic control with chain-of-thought reasoning to enable long-horizon planning.

rt23.webp

RT-2 can use both image and text commands.

Summary

VLMs can be transformed into powerful VLA models to directly control robot by combining VLM pre-training with robotic data.

PaLM-E ∧ PaLI-X → VLA ⟹ improved robotic policies, and significantly better generalisation performance and emergent capabilities inherited from web-scale vision-language pre-training.