Embodied Intelligence Observer

Inferring Action from Future Latent State for Robotic Manipulation

Industry

Source: arXiv cs.ROPublish time unverified

arXiv:2608.22067v1 Announce Type: new Abstract: World-Action Models (WAMs) build robot control on video-generation backbones, which jointly predict dense future visual trajectories and robot actions. We argue that video generation is an unnecessary intermediate objective for world-action modeling.