יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

נש Social Welfare ל Multi Armed Bandits: תיאוריה-מובילה ובטחות גבוהה

Nash Social Welfare for Multi Armed Bandits: Trajectory-wise Expected and High Probability Regret
במאמר זה, המחברים חוקרים פייר מלטי-ארמד באנדיטס תחת המטרה NSW, שמדד את הביצועים דרך המשוקלל גאומטרי של הפרסומים. הם מציגים ניגוד-השקרי-טרג'קטיבי-NSW, שמחשב את המשוקלל הגאומטרי על פני נתיבי-דגימה שלמים, וכן ניגוד-השקרי-NSW-בטחות-גבוהה, שמספק את הניגוד-השקרי הראשון בבטחות-גבוהה בפייר-מלטי-ארמד.
תקציר מקורי באנגליתarXiv:2610.07737v1 Announce Type: cross Abstract: We study fair multi-armed bandits under the Nash Social Welfare (NSW) objective, which measures performance via the geometric mean of accumulated rewards. Existing work defines Nash regret as $\mathrm{NR}_T = \mu^\star - (\prod_{t=1}^T \mathbb{E}\mu_{I_t})^{1/T}$, where $\mu_{I_t}$ is the mean reward of the recommended arm $I_t$ and $T$ is the horizon. Since it applies the geometric mean to per-round marginal expectations, it ignores the joint distribution of rewards across rounds, leaving the NSW fairness motivation unaddressed at the trajectory level. We propose \emph{trajectory-wise Nash regret} $\widetilde{\mathrm{NR}}_T = \mu^\star - \mathbb{E}[(\prod_{t=1}^T \mu_{I_t})^{1/T}]$, which computes the geometric mean over complete sample pa
קרא במקור המקורי