具身智能观察

DeicticVLA:统一语言与指示手势指令模式的 VLA

原标题:DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA

技术动态AI 75

来源:arXiv cs.RO发布时间待核实

arXiv:2608.28108v1 Announce Type: new Abstract: Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category or similar appearance requires detailed expressions that VLAs may not use reliably.

DeicticVLA:统一语言与指示手势指令模式的 VLA | 具身智能观察