Embodied Intelligence Observer

SMILE: Smooth Motion for Improved Long-Horizon VLA Execution

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2608.29432v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models reduce inference cost by executing multiple actions per call, but longer horizons often degrade accuracy because raw chunks contain jitter and outliers. We introduce SMILE, an architecture-preserving interface that predicts B-spline coefficients and decodes them into smooth action sequences.

SMILE: Smooth Motion for Improved Long-Horizon VLA Execution | Embodied Intelligence Observer