כתבה
arXiv cs.LG ·
Range-GRPO: אופטימיזציה של פוליצי על ידי יחסים זוגיים בין תחומי הפרס
Range-GRPO: Policy Optimization via Pairwise Relations among Reward Intervals
Range-GRPO היא פרקטיקה של אופטימיזציה של פוליצי שמשתמשת ביחסים זוגיים בין תחומי הפרס. היא מאפשרת לשפר את הביצועים של LLMs בתחומים שבהם קיימים תחומי פרס.
תקציר מקורי באנגליתarXiv:2610.01548v1 Announce Type: new Abstract: As the use of large language models (LLMs) expands, post-training has become increasingly important for adapting them to downstream tasks. However, obtaining reliable supervision remains costly, especially in domains without reference answers or executable verifiers. LLM-as-a-Judge provides scalable pseudo-rewards for unlabeled responses, but a single point score does not explicitly represent reward uncertainty. This motivates representing pseudo-rewards as conformally calibrated reward ranges. We propose Range-GRPO, a semi-supervised post-training framework that combines limited labeled data with unlabeled prompts. In Group Relative Policy Optimization (GRPO), learning signals depend on relative reward comparisons within each rollout group.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית