Pointing-VLA: Typed Spatial Grounding Interfaces for Vision-Language-Action Manipulation
Industry
Source: arXiv cs.ROPublish time unverified
arXiv:2608.23138v1 Announce Type: new Abstract: Vision-language-action (VLA) models often expose spatial grounding through autoregressive text coordinates or opaque action tokens, creating brittle interfaces between multimodal reasoning and robot execution. We present Pointing-VLA, a typed hidden-state spatial readout built on Embodied-R1. Geometry-specific heads predict normalized points, object-functional grounding (OFG) heatmaps, and visual trajectories without serializing geometry as text.