יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

יתרון תוך-שרשרת: הנחיה של המורה-תלמיד ל-RL עם תגמולים ניתנים לוודאות

Residual Advantage: Student-Relative Teacher Guidance for RL with Verifiable Rewards
המאמר עוסק בפיתוח טכניקה חדשה להנחיה של רשתות עצמאיות, המשתמשת בטכניקה של תגמולים ניתנים לוודאות. הטכניקה, הנקראת Residual Advantage, משתמשת בשיטת תלמיד-מורה כדי להנחות את הרשת, ומציעה תגמולים ניתנים לוודאות לכל צעד בתהליך הלמידה. המאמר מציע תיקון לשיטת ההנחיה הקיימת, ומציע תצוגה חדשה של התגמולים. המאמר כולל תיקון לשיטת ההנחיה הקיימת, ומציע תצוגה חדשה של התגמולים.
תקציר מקורי באנגליתarXiv:2610.11519v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) and on-policy distillation (OPD) have become two main paradigms for post-training reasoning models. RLVR gives each response a single outcome label, leaving the steps inside it without separate credit. OPD provides token-level guidance at student-visited prefixes, but its pointwise signal does not directly reflect the pattern of teacher--student disagreement across the vocabulary. Dense, unbounded log-ratio supervision can amplify the teacher's influence, yet a strong solver is not necessarily a suitable guide when the student's solution paths depart from the teacher's. We propose Residual Advantage (\RA{}), which treats the teacher--student probability residual as a bounded one-step rewar
קרא במקור המקורי