Adaptive Machines - Making Adaptive and Resilient Robots with Generative AI and Reinforcement Learning
Comments
Abstract and Introduction
"Quality-Diversity" algorithms: well known for generating thousands of diverse and high-performing solutions to optimisation task.
Present how Q-D can be paired with Deep RL to learn more complex policies.
OR
how Q-D can be paired with Generative Dynamics Models to ensure fast, safe and continual learning and collection of data.
Application is for cases where environment is dangerous and inaccessible to humans. This also means that there is risk of losing/damaging sensors or actuators on the robot itself.
Traditional damage-recovery approach: self-diagnosis --> pre-designed contingency plan
- Effective when domains and use-cases are simple/straightforward.
- Complexity where adaptive approach is necessitated, becomes difficult to implement.
This paper looks at granular failure/recovery cases.
Intelligent Trial and Error algorithm achieves rapid adaptations combining Q-D with Bayesian Optimisation.
Generating Complex and Diverse Behaviours
Notably, the DCG-MAP-Elites algorithm can generate thousands of complex robotic policies, each with over 20,000 parameters, that solve the same task in unique ways.
Proposed Research Scope
Scenario Generation with LLMs: Leveraging Large Language Models (LLMs) and OMNI-EPIC's environment generation capabilities to create diverse stress-testing scenarios for training the adaptation algorithm. This ensures the algorithm can handle numerous potential issues before deployment.
Hierarchical Adaptation Algorithm: Using the Hierarchical Trial and Error (HTE) algorithm to enhance resilience at all levels, from individual drone failures to fleet-wide reconfiguration and environmental challenges.
Federated Learning for Rapid Adaptation: Adapting Federated Learning methods to enable drones to share their learning across the swarm, speeding up adaptation while identifying and isolating drones with unique mechanical or sensor issues.
Overall Evaluation
2: accept; 1: weak accept; 0: borderline accept; -1: weak reject; -2: reject
2: Accept
Paper Strength
Provide detailed review, including justification of scores.
Starting with the premise, the paper's purpose and intention are clearly defined and valid: adaptivity requires familiarity. And resilience necessitates that the system is secure even when a familiar scenarios is experienced plus some minor variations.
The paper approaches this problem with lens of reinforcement learning. This entails exposing the system to a range of scenarios and letting it evolve over time given a reward function. And rather than relying on real-life occurrences of diverse scenarios, the paper harnesses the power of Generative AI to generate variations of scenarios and subsequent solution-spaces.
The approach is two-pronged (generate solution-space, and train and optimise over the generated solutions). The paper provides evidence towards the efficiency and efficacy of the proposed solution, with a proof-of-concept as support of a six-legged insectoid robot with broken or missing legs.
It is also indicated that the proposed solution can be applied to different systems as well, hinting towards system-agnosticism.
A section dedicated to application of proposed solution to drones and swarms presents a high-level conceptualised-approach. It integrates LLMs, hierarchical adaptation algorithms and federated learning method. The proposed procedure is succinct and supported with a conceptual diagram to illustrate the scope clearly.
The score of 2 is awarded because the premise is clear and relevant. There is prior implementation (meaning that it is not just theoretical) with sufficient documentation and scientific rigour. The methodology is logical and state-of-the-art. Finally, there is a targeted plan of action to adapt the implementation to drones and drone swarms with relevant technologies.
Suggestion to Improve
Provide suggestions that could help improve the paper. Consider commenting on areas such as clarity, methodology, data analysis, relevance of literature reveiw and overall presentation.
Some improvement in presentation required on Figure 1. While meaning can be derived after careful analysis of the figure, it needs to be more descriptive at first glance.
- The graph could be more legible (font and size).
- Some more clarity required on what the box-and-whisker plots represents.
- Rather than the mosaic of pictures on bottom to show the robot's navigation action, a supporting video link that shows this movement in action will have greater impact. The paper provides a video of the LINC research challenge submission, the same could be done for the insectoid robot's adaptation and navigation.
Relevance to SSRC/TII Research
Assess how the submission aligns with the current or future research interests of SSRC and TII.
4: Not Aligned; 3: Neutral; 2: Well Aligned
2: Well Aligned.
The proposed approach to implement the solution with drones and drone swarms aligns with SSRC's current trajectory to explore and secure the UAV and UAV swarm landscape.
It also relies on technologies that are relevant and state-of-the-art, which also aligns with SSRC's approach to explore and augment up an coming technologies.
There is scope for integrating the proposed solution into SSRC's current sub-focus of system recovery to adverse scenarios (along the lines of run-time assurance, system recovery, etc.).
Furthermore, the proposition is system-agnostic to a degree, which makes it potentially applicable to future research interests of SSRC and TII.
Would you suggest this paper for a two year project with SSRC?
Justify answer.
There is scope for adaptation of research into action with drones and drone swarms. Currently, SSRC is transitioning towards securing the drone systems at swarm level. SSRC already has sufficient infrasture in place, meaning that the lead time to first implementation on drone systems can be minimised.
Confidential remarks for the program committee.
If you wish to add any remarks intended only for PC members please write them below. These remarks will only be seen by the PC members having access to reviews for this submission. They will not be sent to the authors. This field is optional.