Embodied Intelligence Observer

Mind-VLA:指令感知的空间表征对齐

Original title: Mind-VLA: Instruction-Aware Spatial Representation Alignment for Vision-Language-Action Models

ResearchAI 80

Source: arXiv cs.ROPublish time unverified

arXiv:2608.04633v2 Announce Type: replace Abstract: Recent Vision-Language-Action (VLA) methods improve generalization by aligning their representations with 3D scene geometry. However, these methods are fundamentally instruction-agnostic: the representations align the entire scene uniformly, neglecting the 3D geometry of the specific target object designated by the language instruction.

Mind-VLA:指令感知的空间表征对齐 | Embodied Intelligence Observer