Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching
技术动态
来源:arXiv cs.RO发布时间待核实
arXiv:2602.00686v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable generalization capabilities in robotic manipulation tasks, yet their substantial computational overhead remains a critical obstacle to real-world deployment. Improving inference efficiency is therefore essential for practical robotic applications.