Embodied Foundation Models (VLA / WAM)
Vision-language-action and world-action models that map vision, language and robot state directly to actions, and reason and plan over long-horizon tasks.
FIR Lab focuses on embodied intelligence: embodied foundation models and learning methods that let robots understand the physical world and carry out complex manipulation reliably in real environments.
Vision-language-action and world-action models that map vision, language and robot state directly to actions, and reason and plan over long-horizon tasks.
Learning how the world changes under robot actions, to predict outcomes, evaluate policies and generate training data.
Learning manipulation skills from demonstrations, teleoperation and egocentric video; shared autonomy, failure recovery and contact-rich manipulation, validated on real robots.
Closed-loop real and synthetic data, simulation for evaluation, and the robot hardware that brings research to real tasks.
People and vision-language-action models collaborating on long-horizon assembly.
The model acts by default and hands control back to the operator when risk rises.
Tactile sensing plus VLA action refinement for contact-rich manipulation.
Multimodal and spatial-temporal reasoning agents coordinating a robot through assembly.
Detecting execution failures in long-horizon tasks and generating recovery strategies.
Evaluating manipulation policies in physics simulation before real deployment.
FIR Lab is part of the Centre for Fundamental and Frontier Sciences, HKISI-CAS, with: