Emergent Collective Intelligence in Swarm Robotics via Stigmergic Multi-Agent Reinforcement Learning
CWSF · 2026 Digital Technology Bronze Medal
Overview
Individual ants possess limited cognitive capability, yet collectively they accomplish complex tasks without central command. This is enabled through stigmergy, an indirect, decentralized coordination via shared environment modification such as pheromone trails. Inspired by my own observations of carpenter ants, this project investigates a novel algorithm for achieving decentralized artificial collective intelligence via stigmergy. A custom simulation was built to train agents through a 17-stage curriculum of increasingly challenging tasks using multi-agent reinforcement learning, coordinating through a shared digital pheromone field. Trained policies were then transferred onto a custom-built physical robot swarm, demonstrating coordinated behaviour in the real world. Experiments confirm that stigmergy is causally beneficial, with pheromone-trained swarms significantly outperforming no-pheromone swarms across all swarm sizes. This open-source platform contributes to Physical AI, with applications in disaster response, planetary exploration, asteroid mining, and ecological monitoring, where resilient, decentralized systems are essential.
Video
This video could not be played here. Watch it on the original project page.
This video could not be played here. Watch it on the original project page.
Why?
Why does artificial intelligence need to be centralized to be powerful?
Problem
Modern AI systems scale intelligence through centralized architectures, increasing model size, computation, and data requirements on specialized hardware [1–4]. In multi-agent swarm systems, as the number of agents grows, complexity rises combinatorially [6]. This makes centralized approaches fundamentally unscalable and inherently fragile.
Inspiration from Nature
Nature offers a fundamentally different model when it comes to swarms. Carpenter ants exhibit sophisticated collective intelligence without central control, despite each individual operating with local sensing, minimal memory, and limited decision-making [7,8]. Coordination emerges through stigmergy: modifying a shared environment by depositing pheromone trails, and responding to those modifications locally [9]. Intelligence is decentralized and emerges collectively from simple agents.
Can artificial collective intelligence emerge from a swarm of simple agents that only sense locally and communicate indirectly through the environment?
Solution
Decentralized artificial collective intelligence offers a more scalable and resilient paradigm [10,11]. Where centralized systems fail at a single point, decentralized swarm systems degrade gracefully [11,12]. Where centralized models require massive infrastructure, decentralized swarm agents operate with minimal individual capacity [10]. This makes stigmergic swarm robotics especially promising for real-world environments where communication is unreliable and centralized control is fragile, such as disaster response, planetary exploration, and beyond [10,13].
How?
Simulation
The simulation is custom-built using PyGame. It provides a 2D environment, modeling agent motion, collisions, obstacles and pheromone fields that diffuse over time, mirroring biological stigmergy [9]. The environment isolates signals from noise to enable stable training.
Reinforcement Learning (RL)
Agents are trained using reinforcement learning to maximize long-term reward through interaction with the environment [14]. This allows complex behaviors to emerge without hard-coding.
Multi-Agent Proximal Policy Optimization (MAPPO)
MAPPO is a multi-agent RL algorithm using a Centralized Training, Decentralized Execution (CTDE) paradigm [15,16]. Agents share a policy updated with Proximal Policy Optimization (PPO) based on overall swarm performance [17]. PPO ensures stable learning by limiting policy changes and preventing destabilizing large updates [17].
During training, a centralized critic with global information evaluates outcomes to guide these updates [16]. During execution, agents act independently using only local observations [15,16].
Gated Recurrent Unit (GRU)
A GRU-based recurrent network is integrated into the policy (and critic) in MAPPO, providing memory [18,19]. This allows agents to retain context and improve coordination [18].
Curriculum Learning
Curriculum learning trains models on tasks of increasing difficulty, allowing it to build foundational skills before tackling complex problems [20]. In this project, training progresses through 17 stages across 3 phases, from simple behaviours to full multi-agent coordination.
This staged progression improves learning stability and enables the emergence of cooperative behaviours [20].
Mission Control
A Mission Control system maintains a global digital pheromone field, receiving deposits and returning relevant pheromone values for each agent. It connects to both simulation and robots through a unified interface and visualizes swarm state and sensor data for real-time monitoring.
Sim-to-Real
The learned policy is transferred to physical robots. On each robot, sensors construct the observation vector, the policy selects an action, and actions are executed directly through motor commands.
What?
Curriculum Learning
The newer 17-stage curriculum-based training configuration, which also added non-carrying outward exploration, nest loiter/crowding penalties, forced outward exploration overrides, post-delivery exit/cooldown handling, and updated promotion thresholds, produced a decisive improvement over an older, simpler recurrent MAPPO training configuration. The newer model achieved 1.25 mean food deliveries per episode and a delivery conversion rate of 58.3%, versus 0.00 deliveries and 0.0% conversion for the older configuration (Welch’s t-test: p = 0.000828, Cohen’s d = 1.254). The older model could pick up food but failed to complete the full foraging loop, showing that the newer curriculum-and-shaping design substantially improved end-to-end foraging performance and was a necessary algorithmic contribution.
Baseline Performance
The MAPPO-based policy significantly outperforms both rule-based and random baselines. The learned policy achieves an average of 1.25 food deliveries, compared to 0.30 for the rule-based method and 0.05 for the random baseline. These improvements are statistically significant (two-sample t-test: p = 0.012250 and p = 0.001235, respectively), confirming the advantage of learned coordination over predefined or stochastic behaviours.
Effect of Stigmergy
Experiments evaluating the role of digital pheromones show that stigmergic communication is critical for coordination. In tasks requiring route reuse, pheromone-enabled swarms significantly outperform those trained without pheromones. At 6 agents, pheromone-enabled policies achieve 1.32 deliveries, while no-pheromone policies fail completely (0.00 deliveries), with statistically significant differences (paired t-test: p = 0.00653). At the new 30-agent upper bound, pheromone-enabled policies achieve 2.66 deliveries, while no-pheromone policies again achieve 0.00, with statistically significant differences (paired t-test: p = 0.000847, Wilcoxon p = 9.96e-06). Late-episode deliveries were also significantly higher at 30 agents (1.12 vs 0.00, paired t-test: p = 0.01868), demonstrating that environmental memory remains an important mechanism for effective coordination at larger swarm sizes.
Swarm Scaling
The scalability of the system was evaluated by measuring exploration coverage and deliveries as a function of swarm size. Results show that performance increases as more agents are introduced, from 0.05 deliveries at 1 agent to 1.25 at 6 agents, and further to 4.65 at 30 agents in the broad final-stage evaluation. This demonstrates that the learned policy enables effective coordination at larger scales without centralized control.
Robustness
The system was tested under challenging conditions including agent loss, sensor noise, and obstacle-dense environments. The swarm demonstrates strong resilience, with only minor performance reductions under agent loss (0.70 deliveries) and sensor noise (1.05 deliveries), compared to the baseline of 1.25. Performance declines more significantly in obstacle-dense environments (0.65 deliveries), identifying a clear limitation and area for future improvement.
Physical Swarm Deployment
The learned policy was successfully deployed onto a swarm of physical robots, demonstrating consistent behaviour between simulation and the real world. The system maintains coordination using a digital pheromone field implemented through Mission Control, validating the transfer of learned behaviours and confirming the feasibility of decentralized swarm intelligence in the real world.
So What?
So What?
This research began with a challenge to a dominant assumption in AI: that intelligence requires large, centralized models [1,2]. Results support an alternative; collective intelligence can emerge from simple, decentralized agents interacting through their environment, without any central controller.
Scientific Impact
This work provides quantitative evidence that stigmergy functions as a scalable coordination mechanism in artificial systems. Environmental memory, rather than direct communication or centralized control, is sufficient to produce emergent collective intelligence. Critically, this coordination scales with swarm size and degrades gracefully under agent failure and sensor noise, confirming the resilience predicted by decentralized system theory [15].
Technical Impact
This research introduces a viable alternative to centralized AI architectures for real-world robotic systems. The combination of curriculum-based MARL, digital stigmergy, and sim-to-real transfer demonstrates that decentralized swarm intelligence is not merely a theoretical concept; it is implementable on physical hardware. The fully open-source platform, including algorithms, models, firmware, mechanical designs, and assembly documentation, makes this research accessible and reproducible.
Applications
Decentralized swarm systems are especially promising where centralized control is impossible and resilience is essential, including disaster response, planetary exploration, asteroid mining, and ecological monitoring. By using the environment as the communication medium, these systems can operate without communication infrastructure, GPS, or a central coordinator.
Conclusion
Intelligence does not need to be centralized to be powerful. This project demonstrates a shift from centralized AI toward decentralized collective intelligence, a paradigm that is scalable, resilient, and grounded in how nature has solved coordination for millions of years.
What's Next?
Future Work
Training with pheromone leads to stronger overall performance, confirming stigmergy as an effective coordination mechanism.
However, pheromone-trained swarms still perform relatively well even when pheromone is removed during evaluation. Results vary by swarm size and coordination phase, suggesting that pheromone may both guide real-time behavior and shape learning during training. Future work will isolate these roles and determine when environmental memory is most critical.
Additionally, ongoing work will evaluate how pheromone-guided behaviors transfer from simulation to physical robots, testing whether learned coordination remains robust under real-world noise, latency, and sensing constraints.
Thanks
Acknowledgements
I would like to thank those who have supported me throughout my science fair journey.
First off, I am extremely grateful for the financial support of my parents and the BC Science Fair Foundation. None of my work would be possible without their help!
I would also like to thank the Greater Vancouver Regional Science Fair committee for their constant hard-work and dedication in supporting me and the rest of the region.
Lastly, I would like to thank Youth Science Canada and the University of Alberta for providing me with the opportunity of a lifetime to connect with my fellow peers across the nation and present my research to an esteemed panel of judges.
References
Research Papers
[1] J. Kaplan et al., “Scaling Laws for Neural Language Models,” arXiv:2001.08361, 2020.
[2] J. Hoffmann et al., “Training Compute-Optimal Large Language Models,” arXiv:2203.15556, 2022.
[3] J. Sevilla et al., “Compute Trends Across Three Eras of Machine Learning,” arXiv:2202.05924, 2022.
[4] R. Bommasani et al., “On the Opportunities and Risks of Foundation Models,” arXiv:2108.07258, 2021.
[5] E. Strubell, A. Ganesh, and A. McCallum, “Energy and Policy Considerations for Deep Learning in NLP,” in Proc. ACL, 2019.
[6] L. Bușoniu, R. Babuška, and B. De Schutter, “A Comprehensive Survey of Multi-Agent Reinforcement Learning,” IEEE Trans. Syst., Man, Cybern. C, vol. 38, no. 2, pp. 156–172, 2008.
[7] T. J. Czaczkes, “Advanced Cognition in Ants,” ResearchGate preprint, 2024.
[8] P. d’Ettorre, P. Meunier, P. Simonelli, and J. Call, “Quantitative cognition in carpenter ants,” Behav. Ecol. Sociobiol., vol. 75, no. 5, 2021, doi: 10.1007/s00265-021-03020-5.
[9] S. Hartman, S. D. Ryan & B. R. Karamched “Walk this way: modeling foraging ant dynamics in multiple food source environments,” Journal of Mathematical Biology, Springer, 2024.
[10] E. Şahin, “Swarm Robotics: From Sources of Inspiration to Domains of Application,” in Swarm Robotics Workshop: State-of-the-art Survey, 2005.
[11] M. Brambilla, E. Ferrante, M. Birattari, and M. Dorigo, “Swarm robotics: a review from the swarm engineering perspective,” Swarm Intelligence, vol. 7, no. 1, 2013.
[12] I. D. Couzin, “Collective cognition in animal groups,” Trends in Cognitive Sciences, vol. 13, no. 1, 2009.
[13] J. Werfel, K. Petersen, and R. Nagpal, “Designing collective behavior in a termite-inspired robot construction team,” Science, vol. 343, no. 6172, 2014.
[14] V. Mnih et al., “Human-level control through deep reinforcement learning,” Nature, vol. 518, no. 7540
[15] T. Yu, J. Chen, S. Gupta, and S. Zhang, “The Surprising Effectiveness of PPO in Cooperative Multi-Agent Games,” arXiv:2103.01955, 2021.
[16] R. Lowe, Y. Wu, A. Tamar, J. Harb, P. Abbeel, and I. Mordatch, “Multi-Agent Actor-Critic for Mixed Cooperative-Competitive Environments,” in Advances in Neural Information Processing Systems (NeurIPS), 2017.
[17] J. Schulman, F. Wolski, P. Dhariwal, A. Radford, and O. Klimov, “Proximal Policy Optimization Algorithms,” arXiv:1707.06347, 2017.
[18] J. Lemmel, R. Grosu, “Real-Time Recurrent Reinforcement Learning,” arXiv preprint arXiv:2311.04830, 2023.
[19] V. Mnih, A. P. Badia, M. Mirza, A. Graves, T. Harley, T. Lillicrap, D. Silver and K. Kavukcuoglu, “Recurrent Neural Networks for Multivariate Time Series with Missing Values,” arXiv preprint arXiv:1606.01865, 2016.
[20] Y. Bengio, J. Louradour, R. Collobert and J. Weston, “Curriculum Learning,” Proceedings of the 26th International Conference on Machine Learning (ICML), 2009, doi:10.1145/1553374.1553380.
[21] M. Rubenstein, C. Ahler, and R. Nagpal, “Kilobot: A low cost scalable robot system for collective behaviors,” IEEE International Conference on Robotics and Automation (ICRA), Saint Paul, MN, USA, May 2012, doi: 10.1109/ICRA.2012.6224638.
[22] M. Dorigo et al., “The e-puck, a Robot Designed for Education in Engineering” ResearchGate preprint, 2009.
[23] F. Arvin, J. Murray, C. Zhang, S. Yue, “Colias: An Autonomous Micro Robot for Swarm Robotic Applications,” International Journal of Advanced Robotic Systems, vol. 11, no. 7, 2014.
Libraries
[24] Python Software Foundation, “Python,” 2026. [Online]. Available: https://www.python.org
[25] NumPy Developers, “NumPy,” 2026. [Online]. Available: https://numpy.org
[26] PyTorch Contributors, “PyTorch,” 2026. [Online]. Available: https://pytorch.org
[27] Pygame Developers, “Pygame,” 2026. [Online]. Available: https://www.pygame.org
[28] Farama Foundation, “Gymnasium,” 2026. [Online]. Available: https://gymnasium.farama.org
[29] J. K. Terry et al., “PettingZoo,” 2026. [Online]. Available: https://www.pettingzoo.ml
[30] Stable-Baselines3 Contributors, “Stable-Baselines3,” 2026. [Online]. Available: https://github.com/DLR-RM/stable-baselines3
[31] Ray Contributors, “RLlib,” 2026. [Online]. Available: https://docs.ray.io/en/latest/rllib
[32] SciPy Developers, “SciPy,” 2026. [Online]. Available: https://scipy.org
[33] PySerial Contributors, “pyserial,” 2026. [Online]. Available:
https://pyserial.readthedocs.io
[34] Bleak Contributors, “bleak,” 2026. [Online]. Available: https://github.com/hbldh/bleak
[35] Bless Contributors, “bless,” 2026. [Online]. Available: https://github.com/kevincar/bless
[36] OpenCV Developers, “OpenCV,” 2026. [Online]. Available: https://opencv.org
Webpages
[37] Raspberry Pi Ltd., “Getting started with Raspberry Pi,” 2026. [Online]. Available: https://www.raspberrypi.com/documentation/computers/getting-started.html
[38] SparkFun Electronics, “How to run a Raspberry Pi program on startup,” [Online]. Available: https://learn.sparkfun.com/tutorials/how-to-run-a-raspberry-pi-program-on-startup/all
[39] OpenAI, “Spinning Up in Deep Reinforcement Learning,” [Online]. Available: https://spinningup.openai.com/en/latest/
[40] Wikipedia, “Swarm robotics,” [Online]. Available: https://en.wikipedia.org/wiki/Swarm_robotics
[41] The Royal Institution, “Engineering a swarm - with Sabine Hauert,” YouTube, 2025. [Online]. Available: https://www.youtube.com/watch?v=E6iJx4ePQCc
Images
[42] Hassell Studio, “Swarm robotics in (interplanetary) onsite construction,” [Online]. Available: https://www.hassellstudio.com/research/swarm-robotics-in-interplanetary-construction. (Image)
[43] CNN, “Minerals are in short supply on Earth. This startup wants to mine asteroids,” [Online]. Available: https://www.cnn.com/world/astroforge-asteroid-mining-nasa-spc-scn. (Image)
[44] The Globe Post, “US Sanctions Can Pose Deadly Obstacles to Humanitarian Aid Delivery,” [Online]. Available: https://theglobepost.com/2018/09/13/aid-sanctions-terrorists/. (Image)
[45] The Ecologist, “How to repel insects and pests naturally,” [Online]. Available: https://theecologist.org/2009/may/01/how-repel-insects-and-pests-naturally. (Image)
[46] Medium, “Reinforcement learning,” [Online]. Available: https://medium.com/analytics-vidhya/reinforcement-learning-4dcd139f82bc. (Image)
[47] Flaticon, “Robot icon,” [Online]. Available: https://www.flaticon.com/free-icon/robot_2083605. (Image)
[48] Wikipedia, “Python logo,” [Online]. Available: https://en.wikipedia.org/wiki/Python_(programming_language)#/media/File:Python-logo-notext.svg. (Image)
[49] Wikipedia, “C++ logo,” [Online]. Available: https://en.wikipedia.org/wiki/C%2B%2B#/media/File:ISO_C++_Logo.svg. (Image)
[50] Amazon, “SainSmart Wide Angle Fish-Eye Camera Lenses for Raspberry Pi 3 Model B Pi 2 Model B+ Arduino, RoHS certified,” [Online]. Available: https://www.amazon.ca/SainSmart-Fish-Eye-Camera-Raspberry-Arduino/dp/B00N1YJKFS. (Image)
[51] Adafruit, “Official Raspberry Pi 5 Active Cooler,” [Online]. Available: https://www.adafruit.com/product/5815. (Image)
[52] Amazon, “DC 5V to 30V Step Up Down Voltage Regulator Module DC-DC Converter Adjustable 1.25V to 30V,” [Online]. Available: https://www.amazon.ca/1-25-30V-Module-Converter-Voltage-Regulator/dp/B07SJLP39M. (Image)
[53] Amazon, “2S Lipo Battery 7.4V 80C 5200mAh with Deans T Plug Rechargeable High Capacity RC Cars Battery Hard Case Fit for RC Car Trucks 1/8 1/10 High Speed RC Cars with Storage Box,” [Online]. Available: https://www.amazon.ca/PCEONAMP-Battery-5200mAh-Rechargeable-Capacity/dp/B0CHJG61NR. (Image)
[54] Amazon, “10Pcs DRV8833 Motor Drive Module 1.5A Dual H Bridge DC Gear Motor Driver Controller Board,” [Online]. Available: https://www.amazon.ca/QCCAN-DRV8833-Module-Bridge-Controller/dp/B0BGLH27GG. (Image)
[55] Zephyr Project, “DOIT ESP32-DevKit-V1,” [Online]. Available: https://docs.zephyrproject.org/latest/boards/others/doit_esp32_devkit_v1/doc/index.html. (Image)
[56] The Pi Hut, “N20 DC Motor with Magnetic Encoder - 6V with 1:100 Gear Ratio,” [Online]. Available: https://thepihut.com/products/adafruit-n20-dc-motor-with-magnetic-encoder-6v-with-1-100-gear-ratio. (Image)
[57] Amazon, “DAOKI 2Pcs VL53L0X Ranging Sensor Time-of-Flight Laser Flight Distance Measurement Sensor Module Breakout 940nm GY-VL53L0XV2 I2C IIC 3.3V/5V,” [Online]. Available: https://www.amazon.ca/DAOKI-Distance-Measurement-Breakout-GY-VL53L0XV2/dp/B08213VH82. (Image)
[58] TinyOS Shop, “SG90 9g micro servo,” [Online]. Available: https://www.tinyosshop.com/sg90-micro-servo. (Image)
[59] National Geographic, “A Swarm of a Thousand Cooperative, Self-organising Robots,” [Online]. Available: https://www.nationalgeographic.com/science/article/a-swarm-of-a-thousand-cooperative-self-organising-robots. (Image)
[60] OpenAI, “AI-generated image of ‘Stigmerman’ using ChatGPT (GPT-5.3), prompt engineered by the author,” 2026.
[61] ResearchGate, “A swarm of e-puck robots equipped with UV-pheromone-module and omni-directional camera releases pheromone in the artificial pheromone environment to accomplish a collective mission,” [Online]. Available: https://www.researchgate.net/figure/A-swarm-of-e-puck-robots-equipped-with-UV-pheromone-module-and-omni-directional-camera_fig3_347711727. (Image)
[62] New Atlas, “Low-cost autonomous robots replicate swarming behavior,” [Online]. Available: https://newatlas.com/colias-swarm-robot/33897/. (Image)
[63] Advanced Science News, “Autonomous robot swarms performing coordinated missions,” [Online]. Available: https://www.advancedsciencenews.com/autonomous-robot-swarms-come-together-to-perform-a-variety-of-missions/. (Image)
Images (30)
Awards (2)
- Bronze Medal
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
ISEF · 2015
The Implantation of the Phototaxic Behavior of Physarum polycephalum in a Decentralized Robotics System
ISEF · 2026
Collective Intelligence: Driving Lessons From Ants
ISEF · 2024
Effective Robotic Swarm Controller Applied to Autonomous Swarm Assembly of Modular Flexible Production Lines
ISEF · 2024
Controlling Robotic Swarms Using LLMs
ISEF · 2024
Autonomous Driving Agents Trained by an Algorithm Based on NeuroEvolution of Augmenting Topologies (Neat) Using a New Logarithmic Fixed-Point Number System Optimized for Microcontrollers
ISEF · 2025
AutoMates: Automating With Self-Learning Intelligent Agents
ISEF · 2025
Chain of Action (CoA): LLM-Powered Multi-Agent Hexapod-Drone System
ISEF · 2025
ASCEND: Autonomous Structure Construction With Novel Agent Swarms
Closest projects by meaning, across every fair and year in the corpus.