כתבה
arXiv cs.LG ·
המטריקה של הסנסור: יתרונות רוחמיים ב-RS-RL תחת פרסומים מעוצבים
The Filter Metric is Safety-Critical: Phantom Advantages in Group-Relative RL under Shaped Rewards
נמצאו יתרונות רוחמיים באופטימיזציה של פוליצי קבוצתית תחת פרסומים מעוצבים. התגלית נעשתה בעזרת חיפוש בארכיון arXiv.
תקציר מקורי באנגליתarXiv:2609.13866v1 Announce Type: new Abstract: Group-relative policy optimization (GRPO and descendants) can discard no-contrast rollout groups through dynamic sampling, while practical implementations expose a configurable filter metric. We identify and quantify a metric-predicate mismatch under composite shaped rewards. When filtering follows the shaped training score rather than the task outcome, all-fail groups retain nonzero within-group spread and pass the predicate; standard-deviation normalization then promotes shaping differences among failures to full-size phantom advantages. In a controlled GSM8K comparison (Qwen2.5-1.5B, LoRA), no filtering and shaped-score filtering end at EM 0.080 +/- 0.112 and 0.040 +/- 0.008, whereas binary-outcome filtering holds 0.754 +/- 0.005 across fo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית