具身智能观察

CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understanding

技术动态

来源:arXiv cs.RO发布时间待核实

arXiv:2608.30289v1 Announce Type: new Abstract: Vision-language-action (VLA) models remain brittle in manipulation tasks that require physical commonsense. Current physical VQA data is typically disembodied and misaligned with robot action domains. Egocentric videos are used only as auxiliary pre-training. It remains unclear whether improved VLM physical understanding actually benefits downstream action generation. Therefore, we present CometVLA to close this gap.

CometVLA: Co-Training on an Embodied Data Pyramid towards Physical Understanding | 具身智能观察