יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

דגימות שכיחות חסרות: תיאור תגמיל פחות מוגדר

Unbiased Reward Modeling from Implicit Feedback for LLM Alignment
הצעה ללמידת תגמיל חסר-מוגדר מפי המשתמש לצורך התאמת LLMs. הצעה זו פותרת את הבעיות של חוסר התגמיל השלילי המוגדר והטיה של תגמיל הפיידבק.
תקציר מקורי באנגליתarXiv:2603.23184v2 Announce Type: replace Abstract: Despite the success of reinforcement learning from human feedback (RLHF), existing reward modeling methods largely rely on explicit feedback, which is costly to collect and difficult to scale. This work studies implicit reward modeling, learning reward models from implicit user feedback, such as clicks, copies and skips. While scalable and cost-effective, implicit feedback poses two key challenges: It lacks definitive negative samples, which makes standard positive-negative classification methods inapplicable; It suffers from selection bias, where responses have heterogeneous propensities to elicit feedback, which further obscures definitive negative samples. To address these challenges, we propose ImplicitRM, which learns unbiased reward
קרא במקור המקורי