From Single-Task Imitation to Real-World Generalization: A Hierarchical Reason-Act Robot System
CWSF · 2026 Digital Technology Silver Medal
Overview
Growing labour shortages and an aging population have created urgent demand for robotic systems capable of assisting across healthcare, elder care, and household settings. Existing approaches rely on single fine-tuned Vision-Language-Action (VLA) models that overfit to their training conditions and fail to generalize to unseen situations. This project develops a hierarchical Reason-Act robot system in which a customized Vision-Language Model (VLM) reasons through and decomposes complex unseen tasks into executable subtasks, while a fine-tuned VLA handles physical execution. A state machine orchestrator coordinates the full system in real time, supported by two continuous improvement pipelines for regression-free performance gains. A voice-driven interface with real-time decision transparency enables natural human-robot interaction. Validated on low-cost 3D-printed hardware, and directly transferable to humanoid and other platforms, the system achieves 97% on single-step tasks and 90% on multi-step unseen tasks, demonstrating that generalization, not memorization, is the path to real-world deployable robotics.
Awards (2)
- Silver Medal
- Selected for CWSF 2026
Competition history
- CWSF 2026
Related projects
ISEF · 2023
Action-Aware Vision Language Navigation: A Novel Reinforcement Learning Framework for Dynamic Navigation in VR and Beyond
ISEF · 2025
EnAct: A Safety-Aware Actionable Guidance System via Multi-Agent Vision-Language Models for People With Vision Impairments
ISEF · 2026
CareBotix in Motion II: Integrating Advanced Manipulation, Autonomous Navigation, and Social Interaction in a Platform Agnostic Robotics System
CWSF · 2026
CareBotix in Motion II: Advanced Robotic Manipulation, Navigation, and Social Interaction
Closest projects by meaning, across every fair and year in the corpus.