כתבה
arXiv cs.AI ·
שיפור הביצועים בלמידת חיזוק
Scoring Higher, Answering Worse: Mitigating Reward Hacking in Rubric-Based RL via Protocol-Level Rubrics
חוקרים הציגו שיטה חדשה לשיפור הביצועים בלמידת חיזוק, Protocol-level Rubrics, שמטרתה למנוע ניצול פרימיות. השיטה מקבצת קריטריונים לממדים ומחייבת את כולם להישאר יחד. התוצאות הראו שיפור של 10.8 נקודות באיכות התשובות.
תקציר מקורי באנגליתarXiv:2609.38847v1 Announce Type: cross Abstract: Rubric-based reinforcement learning (Rubric-RL) trains language models where no verifier exists. A judge checks each criterion of a rubric, and the verdicts are aggregated into a reward, most often by a weighted sum. We show that this additive aggregation is the weak point. Under a sum, criteria compensate for one another: a policy that misses the one decision that matters can buy the points back with advice nobody asked for. On clinical consultation, such a policy scores higher and answers worse. Rubric coverage rises while appropriateness on held-out physician criteria falls below the untrained model. The medical criteria are not to blame. Grouped so that they must hold together, the same criteria, unchanged to the word, recover a third o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית