DeicticVLA:统一语言与指示手势指令模式的 VLA
Original title: DeicticVLA: Unifying Instruction Modes Based on Language and Deictic Gestures in a Single VLA
ResearchAI 75
Source: arXiv cs.ROPublish time unverified
arXiv:2608.28108v1 Announce Type: new Abstract: Vision-Language-Action models (VLAs) allow users to specify manipulation tasks in natural language, but distinguishing a target or placement goal among objects of the same category or similar appearance requires detailed expressions that VLAs may not use reliably.