כתבה
arXiv cs.LG ·
Rising Multi-Armed Bandits with Known Horizons
תקציר מקורי באנגליתarXiv:2602.10727v3 Announce Type: replace Abstract: Rising Multi-Armed Bandits (RMABs) model sequential decision problems where each arm's expected reward improves with repeated pulls. In such problems, the value of investing in an arm depends on how much time remains, making knowledge of the horizon useful side information, yet its benefit remains underexplored. We investigate this benefit through CURE-UCB, a horizon-aware algorithm that estimates each arm's cumulative reward over the remaining horizon. Theoretically, under structured assumptions, we prove that CURE-UCB uniformly dominates a representative horizon-agnostic algorithm and show that the advantage of horizon awareness can be substantial: on some instances, CURE-UCB incurs only $O(1)$ regret whereas the horizon-agnostic algori
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית