Counterfactual rewards promote collective transport using individually controlled swarm microrobots

Veit-Lorenz Heuthe; Emanuele Panizon; Hongri Gu; Clemens Bechinger

doi:10.1126/scirobotics.ado5888

Counterfactual rewards promote collective transport using individually controlled swarm microrobots

Sci Robot. 2024 Dec 18;9(97):eado5888. doi: 10.1126/scirobotics.ado5888. Epub 2024 Dec 18.

Authors

Veit-Lorenz Heuthe^{1

2}, Emanuele Panizon^{3

4}, Hongri Gu¹, Clemens Bechinger^{1

2}

Affiliations

¹ Department of Physics, University of Konstanz, Universitaetsstrasse 10, Konstanz, 78464, Germany.
² Centre for the Advanced Study of Collective Behaviour, Universitaetsstrasse 10, Konstanz, 78464, Germany.
³ Abdus Salam International Centre for Theoretical Physics (ICTP), Strada Costiera 11 Trieste, 34151, Italy.
⁴ Data Engineering Laboratory, Area Science Park, Località Padriciano 99, Trieste, 34149, Italy.

PMID: 39693403
DOI: 10.1126/scirobotics.ado5888

Abstract

Swarm robots offer fascinating opportunities to perform complex tasks beyond the capabilities of individual machines. Just as a swarm of ants collectively moves large objects, similar functions can emerge within a group of robots through individual strategies based on local sensing. However, realizing collective functions with individually controlled microrobots is particularly challenging because of their micrometer size, large number of degrees of freedom, strong thermal noise relative to the propulsion speed, and complex physical coupling between neighboring microrobots. Here, we implemented multiagent reinforcement learning (MARL) to generate a control strategy for up to 200 microrobots whose motions are individually controlled by laser spots. During the learning process, we used so-called counterfactual rewards that automatically assign credit to the individual microrobots, which allows fast and unbiased training. With the help of this efficient reward scheme, swarm microrobots learn to collectively transport a large cargo object to an arbitrary position and orientation, similar to ant swarms. We show that this flexible and versatile swarm robotic system is robust to variations in group size, the presence of malfunctioning units, and environmental noise. In addition, we let the robot swarms manipulate multiple objects simultaneously in a demonstration experiment, highlighting the benefits of distributed control and independent microrobot motion. Control strategies such as ours can potentially enable complex and automated assembly of mobile micromachines, programmable drug delivery capsules, and other advanced lab-on-a-chip applications.