具身智能观察

Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching

技术动态

来源:arXiv cs.RO发布时间待核实

arXiv:2602.00686v2 Announce Type: replace Abstract: Vision-Language-Action (VLA) models have demonstrated remarkable generalization capabilities in robotic manipulation tasks, yet their substantial computational overhead remains a critical obstacle to real-world deployment. Improving inference efficiency is therefore essential for practical robotic applications.

Learning to Accelerate Vision-Language-Action Models through Adaptive Visual Token Caching | 具身智能观察