יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

התאמה פלורליסטית: RLHF מרוב-מפלגתי תחת פידבק אדם-אדם מקוון

Provable Pluralistic Alignment: Multi-Party RLHF under Offline Human Feedback
המאמר עוסק בבעיה של התאמה פלורליסטית בRLHF, כאשר הפידבק האדם-אדם הוא מקוון ופוטנציאלית סתום. החידוש הוא שהמאמר מציע פתרון סטטיסטי סופי לבעיה זו, כולל גם גישה חדשה להתאמה פלורליסטית.
תקציר מקורי באנגליתarXiv:2403.05006v2 Announce Type: replace Abstract: Pluralistic alignment requires learning from feedback that reflects persistent and potentially conflicting stakeholder preferences while ultimately selecting a single collective policy. We study this problem in offline reinforcement learning from human feedback (RLHF), where the party associated with each comparison is observed. Under a shared low-rank linear reward model, we jointly estimate party-specific rewards and perform pessimistic policy optimization under Nash, Utilitarian, and Egalitarian social-welfare objectives. We establish nonasymptotic bounds for party-specific reward estimation and the resulting policy suboptimality under offline coverage conditions. We further consider general pairwise preferences that need not admit a s
קרא במקור המקורי