Embodied Intelligence Observer

超越数据扩展:以表征为中心的 VLA 继续预训练

Original title: Beyond Data Scaling: Representation-Centric Continued Pre-training for Vision-Language-Action Models

ResearchAI 75

Source: arXiv cs.ROPublish time unverified

arXiv:2608.27550v1 Announce Type: new Abstract: Scaling robot data is crucial for building generalist Vision-Language-Action (VLA) models, yet robot trajectories are harder to scale than web-scale image-text data because embodied collection is costly and sparsely covers the physical world.

超越数据扩展:以表征为中心的 VLA 继续预训练 | Embodied Intelligence Observer