יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מינימום חרטה בבנדיטים

Satisficing Regret Minimization in Bandits: Constant Rate and Light-Tailed Distribution
חוקרים פיתחו אלגוריתם חדש למינימום חרטה בבנדיטים. האלגוריתם, SELECT, משיג קצב קבוע של חרטה. נוסף על כך, הם פיתחו גרסה מופחתת, SELECT-LITE, שמשיגה התפלגות קלת-זנב וקצב קבוע של חרטה.
תקציר מקורי באנגליתarXiv:2406.06802v4 Announce Type: replace-cross Abstract: Motivated by the concept of satisficing in decision-making, we consider the problem of satisficing regret minimization in bandit optimization. In this setting, the learner aims at selecting satisficing arms (arms with mean reward exceeding a certain threshold value) as frequently as possible. The performance is measured by satisficing regret, which is the cumulative deficit of the chosen arm's mean reward compared to the threshold. We propose SELECT, a general algorithmic template for Satisficing REgret Minimization via SampLing and LowEr Confidence bound Testing, that attains constant expected satisficing regret for a wide variety of bandit optimization problems in the realizable case (i.e., a satisficing arm exists). As a compleme
קרא במקור המקורי