כתבה
arXiv cs.CL ·
RAWAR: חירוש לפרסומים בלי רול-אאוטס בתחומים ניתנים לוודאות
RAWR: Reward Assignment Without Rollouts in Verifiable Domains
מערכת חדשה לחירוש פרסומים בלי רול-אאוטס, המאפשרת סופרוויזיה גדול-היקף בתחומים ניתנים לוודאות. המערכת, MCNIG, יוצרת סיגנל חזק ואמין, המאפשר חירוש פרסומים זול ומהיר. המערכת נבחנה ב-8 בסימולציות שונות, והציגה תוצאות טובות.
תקציר מקורי באנגליתarXiv:2603.17815v2 Announce Type: replace Abstract: Understanding and evaluating multi-step reasoning in LLMs at the level of individual steps remains a key challenge. Process reward models (PRMs) provide a solution by scoring each step, enabling fine-grained supervision and improved reliability. However, training them requires costly human annotation or computationally intensive rollout-based labeling. To solve this, we introduce MCNIG, a scalable method for automatically labeling the quality of individual reasoning steps in any verifiable domain. Its step score, net information gain (NetIG), improves upon single-reference information gain (IG) by comparing the most-supported correct answer against the most-supported incorrect one, yielding a robust signal even for long and structured out
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית