Embodied Intelligence Observer

Knowing When to Stop: Adaptive Action Chunking via Internal Cross-Attention Dynamics in VLAs

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2609.00908v1 Announce Type: new Abstract: Action chunking is a standard execution strategy in modern Vision-Language-Action (VLA) frameworks, but fixed execution horizons impose a trade-off between efficiency and accuracy. Short chunks require frequent inference and may cause oscillatory behavior, whereas long chunks can become misaligned with newly observed states.

Knowing When to Stop: Adaptive Action Chunking via Internal Cross-Attention Dynamics in VLAs | Embodied Intelligence Observer