כתבה
arXiv cs.LG ·
Principled Direction-Free Intrinsic Motivation through Model-Free Epistemic Free-Energy Estimators
תקציר מקורי באנגליתarXiv:2607.16858v1 Announce Type: new Abstract: Across environments with mixed sources of uncertainty, unsupervised reinforcement learning requires intrinsic motivation that does not precommit to a particular direction of surprise. Surprise minimization is scoped by design to ``unstable'' environments. Prediction-error curiosity rewards total expected surprise, including irreducible noise. Bandit or mixture switching between surprise-minimizing and surprise-maximizing rewards reintroduces non-stationarity by construction. We propose a single intrinsic reward, stationary within each window, derived from the novelty contribution of a preference-free Expected Free Energy objective, expressed in reward-maximization form. Our claim is that parameter information gain, the expected surprise of th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית