GaussVLA: Geometry-Aware Spatial Reasoning for Vision-Language-Action Model
Research
Source: arXiv cs.ROPublish time unverified
arXiv:2608.24959v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models encode visual observations as flat 2D patch tokens that carry no intrinsic geometric structure, and augmenting them with dense monocular depth injects per-pixel scalar values that encode neither surface orientation nor geometric confidence. This leaves the policy with limited structured spatial reasoning for action prediction.