Embodied Intelligence Observer

极简模态掩码即可强化双臂 VLA 鲁棒性

Original title: Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking

IndustryAI 65

Source: arXiv cs.ROPublish time unverified

arXiv:2608.22419v1 Announce Type: new Abstract: Query-based Vision-Language-Action (VLA) models offer low-latency inference that is attractive for bimanual robotic manipulation, but we observe that they can still exhibit discontinuous actions and execution failures in complex dual-arm tasks. We hypothesize that unstable multi-view and language fusion is one contributing factor in these failures, often coinciding with attention spreading to distracting regions.

极简模态掩码即可强化双臂 VLA 鲁棒性 | Embodied Intelligence Observer