具身智能观察

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

技术动态

来源:arXiv cs.RO发布时间待核实

arXiv:2608.07267v2 Announce Type: replace-cross Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions. Although semantically capable, such action-centric training does not explicitly model how the agent's visual observations should evolve under its predicted motion.

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN | 具身智能观察