Embodied Intelligence Observer

CLAP:跨本体视频世界模型即零样本物理仿真器

Original title: CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

IndustryAI 82

Source: arXiv cs.ROPublish time unverified

arXiv:2608.27406v1 Announce Type: new Abstract: State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents.

CLAP:跨本体视频世界模型即零样本物理仿真器 | Embodied Intelligence Observer