具身智能观察

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

技术动态

来源:arXiv cs.RO发布时间待核实

arXiv:2508.13073v3 Announce Type: replace Abstract: Robotic manipulation, a key frontier in robotics and embodied AI, requires precise motor control and multimodal understanding, yet traditional rule-based methods fail to scale or generalize in unstructured, novel environments. In recent years, Vision-Language-Action (VLA) models, built upon Large Vision-Language Models (VLMs) pretrained on vast image-text datasets, have emerged as a transformative paradigm.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey | 具身智能观察