具身智能观察

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

技术动态

来源:arXiv cs.RO发布时间待核实

arXiv:2608.30935v1 Announce Type: new Abstract: Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embodiments. Modern vision-language models (VLMs) already encode spatial priors for visual grounding, spatial reasoning, and pointing, but these capabilities are rarely elicited directly for robot control.

多源报道2

查看事件全景 →
LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation | 具身智能观察