כתבה
arXiv cs.LG ·
When Greedy Sampling Explores: KL-Regularized Contextual Bandits without Eluder-Dimension Dependence
תקציר מקורי באנגליתarXiv:2609.13564v1 Announce Type: new Abstract: We study KL-regularized contextual bandits under both reward and preference feedback. We show that greedy sampling can achieve logarithmic regret without explicit dependence on the eluder dimension. For reward feedback, we establish an eluder-dimension-independent regret bound for a simple greedy algorithm that directly samples from the Gibbs policy induced by the estimated reward. We further extend this result to preference feedback under both the general preference and Bradley--Terry models, while also sharpening existing dimension-dependent guarantees. Our analysis reveals a trade-off between greedy sampling and upper confidence bound-style exploration: greedy sampling enjoys stronger guarantees when KL regularization is sufficiently stron
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית