יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אינפלציה של תגמול: גירוי בריא ללמידת חיזוק

Reward Inflation: A Healthy Stimulus for Reinforcement Learning
חוקרים הציעו 'אינפלציה של תגמול' - שיטה לשיפור למידת חיזוק. השיטה מעלה את רמת התגמול במהלך האימון, מה שמאפשר הסתגלות מהירה יותר. התוצאות הראו שיפור בביצועים במשימות שונות.
תקציר מקורי באנגליתarXiv:2610.02545v1 Announce Type: new Abstract: Reward serves as the primary learning signal in reinforcement learning (RL). However, while reward magnitudes are typically held fixed throughout training, their temporal modulation remains underexplored. In this paper, we propose reward inflation, a gradual scaling of rewards over the course of training, and show that it can act as a healthy stimulus for RL. Theoretically, reward inflation induces an implicit recency weighting that upweights recent transitions during policy updates, enabling faster adaptation. We further show that, by sustaining gradient signals as the policy saturates, reward inflation suppresses the emergence of dormant neurons and helps preserve plasticity. Empirical results on ALE games and MuJoCo tasks corroborate these
קרא במקור המקורי