כתבה
arXiv cs.LG ·
Learning from World Feedback: Why Model Uncertainty Fails as a Risk Signal in Model-Based RL
תקציר מקורי באנגליתarXiv:2607.16591v1 Announce Type: new Abstract: The RLxF programme argues that learning signals should come from world feedback rather than from internal model proxies. We instantiate this position in safe model-based control and distil it into three concrete design principles. Empirically, across four world-model architectures spanning a 2x MSE range, MPC planning is statistically equivalent (TOST, n=200), and dynamics-based uncertainty penalties increase collision rates from 26% to 34%: the standard MBRL safety proxy is anti-correlated with safety in this regime. Replacing the model-internal proxy with three world-feedback signals (a sensor-derived margin via minimum lidar, a temporal signal via time-to-collision, and an outcome-supervised feedback model g_psi trained on prior collision
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית