Embodied Intelligence Observer

RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2512.23649v5 Announce Type: replace Abstract: Humans learn locomotion through visual observation, interpreting visual content first before imitating actions. However, state-of-the-art humanoid locomotion systems rely on either curated motion capture trajectories or sparse text commands, leaving a critical gap between visual understanding and control.

RoboMirror: Understand Before You Imitate for Video to Humanoid Locomotion | Embodied Intelligence Observer