כתבה
arXiv cs.LG ·
Allocation Stability and Wald Inference under Variance-Aware UCB
תקציר מקורי באנגליתarXiv:2412.08843v3 Announce Type: replace-cross Abstract: Allocation stability is often used to justify Gaussian inference from bandit data, but when is it necessary? In this paper, we address this question for a two-armed, fixed-horizon variance-aware UCB policy with bounded reward distributions that may vary with the horizon. We find a sharp criterion in terms of the reward gap and variances that determines whether the optimal-arm count admits a deterministic approximation with vanishing relative error, while the suboptimal-arm count is always stable. Despite the possible instability of the optimal-arm count, we show that the ordinary Wald statistic for a linear combination of the arm means has a standard normal limit for every fixed nonzero coefficient vector, provided the product of th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית