כתבה
arXiv cs.CL ·
מה מערכות רווארד מזכירות?
What do Reward Models Memorize?
מחקר זה בוחן מה מערכות רווארד מזכירות. התוצאות מראות שמערכות רווארד מזכירות זוגות העדפה קלים, קיצורי דרך והיבטים פשוטים. הממצאים מצביעים על כך שמערכות רווארד מתוכננות באופן מוטה.
תקציר מקורי באנגליתarXiv:2607.24484v1 Announce Type: cross Abstract: This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.g., model identity, user sampling strategy), and 3) overgeneralize simple heuristic correlates of human preference (e.g., length, compliance) when confronted with unseen preference pairs. Overall, our findings indicate that discriminative training of RMs from human preference data results in biased RMs not yet capable of judging response quality in context-dependent scenarios.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית