Embodied Intelligence Observer

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2508.13073v3 Announce Type: replace Abstract: Robotic manipulation, a key frontier in robotics and embodied AI, requires precise motor control and multimodal understanding, yet traditional rule-based methods fail to scale or generalize in unstructured, novel environments. In recent years, Vision-Language-Action (VLA) models, built upon Large Vision-Language Models (VLMs) pretrained on vast image-text datasets, have emerged as a transformative paradigm.

Large VLM-based Vision-Language-Action Models for Robotic Manipulation: A Survey | Embodied Intelligence Observer