יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Range-GRPO: אופטימיזציה של פוליצי על ידי יחסים זוגיים בין תחומי הפרס

Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals
Range-GRPO היא פרקטיקה של אופטימיזציה של פוליצי שמשתמשת ביחסים זוגיים בין תחומי הפרס. היא מאפשרת לשפר את הביצועים של LLMs בתחומים שבהם קיימים תחומי פרס.
תקציר מקורי באנגליתarXiv:2610.01548v1 Announce Type: new Abstract: As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group.
קרא במקור המקורי