יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

STAR-GRPO: תקן-אנכורינג ויתרונות-בטוחים נגד חקירה-שכירה של תיאור-התייצוג

STAR-GRPO: Canonical Anchoring and Reliability-First Advantages against Representation-Dependent Reward Hacking
מערכת STAR-GRPO מציעה פתרון להטרדה-שכירה באופטימיזציה של פוליצי-קבוצתית. המערכת מספקת ערך-יתרון ראשון-בטוח, המבוסס על ניתוח-זוגי של אותו-פעל-מעקב. STAR-GRPO נבחן בשני סביבות-הכשרה שונות. בסביבת-הכשרה של חקירה-שכירה באופן-טוקן, STAR-GRPO מקטין את האופטימיזציה של הסקור-המוטמע. בסביבת-הכשרה של חקירה-שכירה באופן-רוביק, STAR-GRPO משפר את הביקורת-עצמאית, מצמצם את הפער-השפה, ומקטין את הטענה-המוגזמת.
תקציר מקורי באנגליתarXiv:2609.36900v1 Announce Type: new Abstract: Reward hacking occurs when policy optimization exploits a brittle reward interface or an overly permissive proxy objective, improving the training score without improving the underlying response quality. This phenomenon is amplified in group-relative policy optimization: an unsupported reward can shift the group baseline and alter the updates of other rollouts, while post-hoc or purely relative weighting cannot represent group-wide uncertainty. We propose \emph{Self-Tuned Anchored Reliability Group-Relative Policy Optimization} (STAR-GRPO), a reliability-first advantage estimator based on paired assessments of the same rollout. STAR separates the quality signal from its learning influence: score disagreement determines rollout reliability, re
קרא במקור המקורי