Learning to Act While Waiting: RL Finetuning of Generalist Robot Policies Under Inference Latency
Industry
Source: arXiv cs.ROPublish time unverified
arXiv:2608.23831v2 Announce Type: replace Abstract: While reinforcement learning (RL) allows generalist robot policies to continually improve during deployment, the large model size of modern generalist policies, such as VLAs, poses a fundamental obstacle to effective RL improvement.