כתבה
arXiv cs.LG ·
Learning a Ranking from Human Feedback in Log-Concave Random Utility Models
תקציר מקורי באנגליתarXiv:2610.07973v1 Announce Type: new Abstract: We study the problem of recovering the ranking of a fixed set of items according to their unknown numerical utilities. At each interaction with the environment, a learner presents the item set to a human and receives comparative feedback of two types. Under full-ranking feedback, each interaction reveals a noisy ranking of all items, whereas under winner-only feedback, it reveals only the item ranked first. In both settings, we model human feedback using a random utility model with log-concave noise and study the number of observations needed to recover an $\epsilon$-accurate ranking with high probability. This novel criterion tolerates ordering errors only between items whose utilities differ by less than $\epsilon$. For both feedback types,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית