כתבה
arXiv cs.LG ·
Range-GRPO: אופטימיזציה של מדיניות דרך יחסים זוגיים בין אינטרוולים של תגמול
Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals
Range-GRPO הוא כלי חדש לאופטימיזציה של מדיניות, המשתמש ביחסים זוגיים בין אינטרוולים של תגמול. הוא מאפשר לשפר את ביצועי המודלים במשימות שונות, תוך שימוש בנתונים מוגבלים. Range-GRPO מתאים למשימות עם אינטרוולים של תגמול, ויכול לשפר את הביצועים בהשוואה לשיטות אחרות.
תקציר מקורי באנגליתarXiv:2610.01548v2 Announce Type: replace Abstract: As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout gro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית