כתבה
arXiv cs.LG ·
On the Complexity of Preference-Based Bandits
תקציר מקורי באנגליתarXiv:2609.39351v1 Announce Type: cross Abstract: We study preference-based bandits with general reward function classes, where a learner sequentially selects pairs of arms and observes binary preference feedback governed by the Bradley--Terry model. This setting naturally arises in applications such as recommender systems, tournament ranking, and learning from human feedback, where relative preferences are easier to elicit than absolute rewards. The observation model inherits the logistic bandit challenge of handling the problem-dependent constant $\kappa$, which accounts for the non-linearity of the link function and can grow arbitrarily large. Moreover, prior work has predominantly focused on linear or kernelized reward models, precluding the use of richer function classes. To address t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית