יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

RankBuffer: שיטה חדשה לתגמול יעיל

RankBuffer: Efficient Ranking-Based Rewards for Open-Ended Generation
RankBuffer היא שיטה חדשה לתגמול יעיל בייצור פתוח. היא משתמשת ברשימה מסודרת של תגובות קודמות כדי לחסוך בעלויות שיפוט. השיטה הוכחה כיעילה בארבעה מבחנים שונים.
תקציר מקורי באנגליתarXiv:2609.36652v1 Announce Type: new Abstract: Open-ended generation lacks canonical answers, making pointwise rewards difficult to calibrate for group-based reinforcement learning. Directly ranking same-query rollouts provides a more suitable relative reward signal, but existing ranking-based reward methods can incur substantial judging cost. We introduce RankBuffer, which maintains an ordered, query-specific buffer of previously judged responses as a reusable quality scale. Each rollout is first inserted into an anchor interval through an independent coarse judgment, after which only rollouts assigned to the same interval undergo local fine ranking. The resulting complete order is converted into bounded rank rewards, while boundary expansion, local refinement, and inactive-anchor prunin
קרא במקור המקורי