具身智能观察

Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models

技术动态

来源:arXiv cs.RO发布时间待核实

arXiv:2608.30643v1 Announce Type: new Abstract: Recent vision-language-action (VLA) methods improve manipulation performance by aligning their representations with 3D scene geometry. However, these methods often struggle with long-horizon manipulation and observation aliasing between visually similar states due to a lack of temporal information: the 3D scene geometry captures only the current state, rather than how it has evolved over time.

Temporal Forcing: 4D Representation Alignment for Vision-Language-Action Models | 具身智能观察