Embodied Intelligence Observer

GeoWAM:面向自动驾驶的视觉几何世界动作模型

Original title: GeoWAM: Visual Geometry World Action Models for Autonomous Driving

IndustryAI 65

Source: arXiv cs.ROPublish time unverified

arXiv:2608.23486v1 Announce Type: cross Abstract: World action models (WAMs) have recently gained increasing attention as a framework for jointly modeling scene evolution and ego actions in autonomous driving. Most existing WAMs learn scene dynamics in pixel space by combining a video-generation backbone for future-observation prediction with an action head for ego-trajectory prediction.

GeoWAM:面向自动驾驶的视觉几何世界动作模型 | Embodied Intelligence Observer