כתבה
arXiv cs.LG ·
Multi-User Dueling Bandits: A Fair Approach using Nash Social Welfare
תקציר מקורי באנגליתarXiv:2605.01961v3 Announce Type: replace Abstract: Learning from human preference data is becoming a useful tool, from fine-tuning large language models to training reinforcement learning agents. However, in most scenarios, the model is trained on the average preference of all human evaluators, which, under large variations of preferences, can be unfair to minority groups. In this work, we consider fairness in dueling bandits, a standard framework for online learning from preference data. We assume that each user has a (potentially distinct) Condorcet winner, which is an arm preferred to every other arm. Using these user-specific Condorcet winners as reference points, we evaluate and score arms according to their performance relative to the corresponding winner. To promote fairness across
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית