具身智能观察

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning

产业动态

来源:arXiv cs.RO发布时间待核实

arXiv:2608.22591v1 Announce Type: new Abstract: Robot policies receive heterogeneous observations at each decision step, yet sequence models differ in how they organize these inputs over time. We introduce WorldToken, a time-first policy instantiation that fuses multiview images, proprioception, and task conditioning within each policy timestep into one world token. A causal temporal Transformer models the resulting world-token sequence, and a diffusion action head generates action chunks.

WorldToken: Time-First Sequence Modeling for Robotic Imitation Learning | 具身智能观察