יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Range-GRPO: אופטימיזציה של מדיניות דרך יחסים זוגיים בין אינטרוולים של תגמול

Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals
Range-GRPO הוא כלי חדש לאופטימיזציה של מדיניות, המשתמש ביחסים זוגיים בין אינטרוולים של תגמול. הוא מאפשר לשפר את ביצועי המודלים במשימות שונות, תוך שימוש בנתונים מוגבלים. Range-GRPO מתאים למשימות עם אינטרוולים של תגמול, ויכול לשפר את הביצועים בהשוואה לשיטות אחרות.
תקציר מקורי באנגליתarXiv:2610.01548v2 Announce Type: replace Abstract: As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout gro
קרא במקור המקורי