כתבה
arXiv cs.CL ·
כשהצלחה בתחום ההפצה נכשלת: בחינת דגמי תמריצים חלש-חזק תחת שינוי טיב העדות
When In-Distribution Gains Fail: Evaluating Weak-to-Strong Reward Models under Preference Shift
במאמר זה, נבחן דגמי תמריצים חלש-חזק תחת שינוי טיב העדות. נמצא כי דגמי תמריצים חזקים שהוכשרו על תמריצים חלשים יכולים להצליח בתחום ההפצה, אך נכשלים בהעברה לתחומים אחרים. נציג פתרון חדש, Representation Anchoring, שמטרתו למנוע דחיפה יתרה של הדגם כלפי תחום ההפצה המקורי.
תקציר מקורי באנגליתarXiv:2605.25629v3 Announce Type: replace Abstract: Weak-to-strong (W2S) generalization is a promising framework for scalable oversight, yet existing evaluations often test students under matched train-test distributions. Therefore, we study W2S preference learning under zero-shot distribution shift and find that strong students trained on weak preference labels can appear successful in-distribution while failing to transfer across preference datasets. We provide evidence for a representational failure mode in which weak-supervised fine-tuning can pull the strong model toward source-domain features instead of maintaining broadly transferable preference representations. To mitigate this, we propose Representation Anchoring (Anchor), a simple yet effective regularizer that constrains excessi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית