Mind-VLA:指令感知的空间表征对齐
Original title: Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models
ResearchAI 80
Source: arXiv cs.ROPublish time unverified
arXiv:2608.04633v2 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are fundamentally instruction-agnostic: the representations align the entire scene uniformly, neglecting the 3D geometry of the specific target object designated by the language instruction.