TrAct:用视觉轨迹桥接机器人控制与视觉预测
Original title: TrAct: Bridging Robot Control and Visual Prediction with Visual Tracks
ResearchAI 85
Source: arXiv cs.ROPublish time unverified
arXiv:2608.24101v3 Announce Type: replace Abstract: Robot actions are inherently embodiment-specific and only weakly aligned with image-space visual changes, limiting their effectiveness as conditioning signals for robot world models. In contrast, visual tracks provide an embodiment-agnostic representation of how task-relevant points move through a scene, offering dense image-space guidance for accurate and spatially precise future video prediction.