כתבה
arXiv cs.LG ·
הטוב משני העולמות לבנדיטים רציפים עם תצפיות משולמות
Best-of-Both Worlds for linear contextual bandits with paid observations
אנו מחקרים בנבדיטים רציפים עם תצפיות משולמות, ומציגים שני אלגוריתמים שמבטיחים טוב-משני-העולמות. האלגוריתם Agg-SPB משלב יציבות תלויה במצב, והאלגוריתם CE-SPB משלב סבירות תלויות בארמים. שני האלגוריתמים מגיעים לתוצאות טובות בסביבות סטוכסטיות ובסביבות קוראוטיות.
תקציר מקורי באנגליתarXiv:2510.07424v3 Announce Type: replace Abstract: We study linear contextual bandits with paid observations, where at each round the learner observes a context, selects an action, and may pay a fixed cost to observe feedback from a subset of arms. We propose two Follow-the-Regularized-Leader algorithms with Best-of-Both-Worlds guarantees. The first, Agg-SPB, extends the SPB-matching framework of Tsuchiya and Ito (2024) by aggregating context-dependent stability terms, achieving the characteristic $T^{2/3}$ adversarial regret rate and logarithmic dependence on $T$ in stochastic environments. The second, CE-SPB, combines arm-dependent observation probabilities with an entropy-adaptive learning rate inspired by Kuroki et al. (2024). It achieves an entropy-adaptive $\widetilde{O}(T^{2/3})$ a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית