Embodied Intelligence Observer

RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation

Research

Source: arXiv cs.ROPublish time unverified

arXiv:2608.25585v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models provide a versatile foundation for general robotic manipulation, yet they exhibit significant brittleness when confronted with novel task distributions. While In-Context Imitation Learning (ICIL) offers a training-free alternative, existing frameworks suffer from an adaptation bottleneck that hinders the effective translation of expert context to executable actions.

RA-VLA: Retrieval-Augmented VLA for Test-Time Adaptation | Embodied Intelligence Observer