具身智能观察

策略感知模拟器学习的理论基础与高效算法

原标题:Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning

技术动态AI 65

来源:arXiv cs.LG发布时间待核实

arXiv:2605.29032v3 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably exploit minor model inaccuracies, leading to simulator exploitation and a reality gap where policies succeed in simulation but fail in the real world.

策略感知模拟器学习的理论基础与高效算法 | 具身智能观察