כתבה
arXiv cs.AI ·
למידת השלכה מרוב-אורך-זמן
Adaptive Multi-Horizon Reinforcement Learning
אורח חיים של למידת השלכה, המתאים לסביבות מורכבות. המאמר עוסק בלמידת השלכה מרוב-אורך-זמן, שמאפשרת לסקור ולהתאים לשינויים במבנה הפרס.
תקציר מקורי באנגליתarXiv:2607.20656v1 Announce Type: cross Abstract: Effective decision-making in complex and changing environments requires balancing short-term and long-term consequences. In reinforcement learning (RL), this trade-off is typically controlled through a fixed discount factor, which imposes a single exponentially discounted temporal horizon. However, biological agents exhibit flexible and adaptive temporal discounting, suggesting that effective planning requires multiple timescales. Here, we propose a multi-horizon approach that adaptively selects and combines temporal horizons, enabling robust adaptation to changes in reward structure without manual discount-factor tuning. This flexibility makes the method particularly suitable for continual learning scenarios involving task switches and var
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית