具身智能观察

LD4WAM:从人类视频学习隐动力学供世界动作模型

原标题:LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

产业动态AI 69

来源:arXiv cs.RO发布时间待核实

arXiv:2608.22403v1 Announce Type: new Abstract: Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. However, most WAMs learn from such video only by predicting pixel-level future frames, giving dynamics that are not directly actionable, whereas motion retargeting recovers directly actionable actions but leaves a large visual gap across embodiments.