כתבה
arXiv cs.LG ·
מתי פרסומים פנימיים מובילים לחקירה?
When Do Intrinsic Rewards Lead to Exploration?
במאמר זה נבחן את השאלה מתי פרסומים פנימיים מובילים לחקירה. המחברים מציעים קריטריון רשמי לחקירה שמשווה מדיניות על ידי המידע הפוטנציאלי שהן יכולות לצבור.
תקציר מקורי באנגליתarXiv:2610.02159v1 Announce Type: new Abstract: Intrinsic rewards are designed to guide exploration in reinforcement learning by assigning value to an agent's experience, for example through prediction error or learning progress. However, maximizing these rewards need not produce the most informative experience available. We propose a formal criterion for exploration that compares policies by the counterfactual information they acquire: how well their histories can substitute for experience under alternative policies. We construct a single, simple environment in which specified count-based, prediction-error, empowerment, and information-gain objectives have maximizing policies that are Pareto-suboptimal at acquiring counterfactual information. We explain these failures and establish condit
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית