Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models
Research
Source: arXiv cs.ROPublish time unverified
arXiv:2608.30643v1 Announce Type: new Abstract: Recent vision-language-action (VLA) methods improve manipulation performance by aligning their representations with 3D scene geometry. However, these methods often struggle with long-horizon manipulation and observation aliasing between visually similar states due to a lack of temporal information: the 3D scene geometry captures only the current state, rather than how it has evolved over time.