策略感知模拟器学习的理论基础与高效算法
原标题:Theoretical Foundations and Effective Algorithms for Policy-Aware Simulator Learning
技术动态AI 65
来源:arXiv cs.LG发布时间待核实
arXiv:2605.29032v3 Announce Type: replace Abstract: Model-based reinforcement learning (MBRL) agents typically learn world models by minimizing predictive loss. However, powerful RL optimizers inevitably exploit minor model inaccuracies, leading to simulator exploitation and a reality gap where policies succeed in simulation but fail in the real world.