Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
Research
Source: arXiv cs.ROPublish time unverified
arXiv:2602.00686v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable generalization capabilities in robotic manipulation tasks, yet their substantial computational overhead remains a critical obstacle to real-world deployment. Improving inference efficiency is therefore essential for practical robotic applications.