Embodied Intelligence Observer

Flex-$\pi$: A Multi-Stream World-Action Model with Compute Flexibility

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2608.10860v3 Announce Type: replace Abstract: World-action models (WAMs) predict the future to act better, but nearly all of them predict only RGB latents, trained purely for pixel reconstruction, with no explicit signal for the 3D geometry or object semantics manipulation needs. We find a surprising free lunch: the same frozen video-generation VAE that encodes RGB also encodes 3D pointmaps almost losslessly, with no pointmap-specific training at all.