Embodied Intelligence Observer

TFP: Temporally Conditioned Memory-Fusion Policies for Visuomotor Learning

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2607.08283v3 Announce Type: replace Abstract: Vision--Language--Action (VLA) policies such as $\pi_{0.5}$ and OpenVLA perform well on many manipulation tasks, but they are often reactive: the next action is predicted from the current observation, instruction, and proprioceptive state. This assumption breaks down in stage-dependent manipulation, where visually similar states may require different actions depending on latent task progress and previous interaction outcomes.

TFP: Temporally Conditioned Memory-Fusion Policies for Visuomotor Learning | Embodied Intelligence Observer