Enabling Novel Mission Operations and Interactions with ROSA: The Robot Operating System Agent
Abstract
- ROSA is an AI-powered agent, that bridges the gap between ROS and natural language interfaces.
- Using SoTA language models with Open Source framworks
- ROSA enables operators to interact with robots using natural language
- Translates commands into actions and interfacing with ROS through well-defined tools
- Integrates with both ROS 1 and ROS 2 through its modular design.
- Can also be extended to interface with other robot middleware.
Introduction
- ROSA leverages large language models and open-source software like ROS and LangChain to implement Reasoning and Acting (ReAct) agent that can understand and execute commands on its robotic host.
Aim
Make Robotic systems more accessible and userfriendly, empowering a broader range of operators to interact with robots in a more intuitive manner.
- Target Audience
- Designed to benefit wide range of users:
- Robotics researchers and developers - Streamline development, verification and validation.
- Field operators and technicians - Perform routine operations without needing specialised training and understanding of ROS.
- Educators and students - Enhance learning experience without prerequisite of Linux, CLI, and ROS expertise.
- Hobbyist and makers - Engage with robotics in a more accessible manner.
- Rapid deployment and close to metal control with an LLM abstraction layer above it.
- Designed to benefit wide range of users:
Background
- ROSA constitutes an Embodied Agent.
- ROSA adopts ReAct agent paradigm:
- Tools developed for ROSA act as wrappers around standard ROS and ROS2 utilities like
rosnode,rostopic, andrviz.
Agent Architecture
-
Action Space
- Set of tools agent can call within working environment.
- Subset of standard ROS tools.
- Several utility commands and functions ROSA can invoke in response to user queries.
- Standard ROS tools included.
- Developers can extend ROSA's capacity by adding custom functions tailored to specific robotic platforms or applications.
- ROSA incorporates mechanisms to prevent invocation of unsafe or unauthorised actions.
- [?] How does it do this? -> Safety and Constraint Handling
- Set of tools agent can call within working environment.
-
Robot System Prompts (RSP)
- Set of prompts to help guide ROSA's behaviour.
- Supply LLM with essential insights about:
- robot identity
- environment
- operating conditions
- Purpose:
- Understand how robot presents itself to user and interacts with human operators.
- Establish clear operational boundaries (safety protocols and limitations).
- Anticipate and react to environmental changes in real-time.
-
Tool invocation and Multi-Tool usage
- Rosa interprets user's NL request and identifies appropriate tool(s) to invoke to fulfil the request.
- Process:
- Forward query to LLM w/ list of available tools, RSP and chat history.
- LLM interprets intent of query and maps to it one+ tools within ROSA's action space.
- LLM also determines necessary parameters for the tool based on user request (fills defaults and asks for follow up if needed).
- Action invoked within ROS environment.
- ROSA forwards response of tool back to LLM to observe out of action to generate feedback.
-
Safety and Constraint Handling
- Before invoking commands, ROSA validates parameters to ensure they are within acceptable ranges.
- Certain actions, services, parameters, or topics can be blacklisted.
- Critical actions can request for permission from operator.
- [!] This doesn't appear totally "Intelligent" in my opinion.
-
Custom Agent
- Specialisation of agent to specific application.
- Custom prompts and tools to enforce safety protocols specific to the application.
- Tailor persona and interaction style to suit operator preference.
- Extensibility to allow for addition of new functionalities.
Implementation
@tooldecorator from LangChain library used on functions to register them as actions the language model can invoke.- ROSA tools return structured data in the form of dictionaries or lists.
- By providing more functions (tools) can limit the LLM's ability to hallucinate by providing the Agent means to run those calculations
Human-Robot Interaction
- Bridging knowledge gap with AI assistance
- Multimodal interaction capabilities
- Enhancing communication and collaboration
- Supporting diverse operator expertise levels
Ethics for Embodied Agents
Some ethics BS incorporating Asimov's Three Laws and what not.
Personally, this is all just hocus pocus that they've written to pad the content.
