极简模态掩码即可强化双臂 VLA 鲁棒性
原标题:Robust Bimanual Vision-Language-Action Models via Embarrassingly Simple Modality Masking
产业动态AI 65
来源:arXiv cs.RO发布时间待核实
arXiv:2608.22419v1 Announce Type: new Abstract: Query-based Vision-Language-Action (VLA) models offer low-latency inference that is attractive for bimanual robotic manipulation, but we observe that they can still exhibit discontinuous actions and execution failures in complex dual-arm tasks. We hypothesize that unstable multi-view and language fusion is one contributing factor in these failures, often coinciding with attention spreading to distracting regions.