כתבה
arXiv cs.LG ·
הפרשה של התאמת AI: RLHF היא תאמן יוצרני סביל
Distortion of AI Alignment Revisited: RLHF is a Decent Utilitarian Aligner
בעיית ההפרשה של RLHF בסביבות פלורליסטיות נשאלה, אך בדיקה פילוסופית מציגה שאינה תכונה יסודית של האלגוריתם.
תקציר מקורי באנגליתarXiv:2609.12651v1 Announce Type: new Abstract: While Reinforcement Learning from Human Feedback (RLHF) is the standard paradigm for aligning large language models with human preferences, its effectiveness in pluralistic settings has been called into question. Notably, recent work by G\"olz et al. (2025) demonstrated that the \textit{distortion} -- defined as the multiplicative gap between the average user utility of the RLHF policy and the optimal average utility -- can scale exponentially with the Bradley-Terry temperature parameter $\beta$ when users have heterogeneous preferences. In this work, we present a fine-grained analysis of the distortion of RLHF with reward clipping and demonstrate that such exponential degradation is not a fundamental property of the algorithm but rather a co
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית