יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הפרדת חקירה מאופטימיזציה ב-RLVR

Decoupling Exploration from Optimization in RLVR
במאמר זה, המחברים מציגים פרקטיקה חדשה של הפרדת חקירה מאופטימיזציה ב-RLVR, שמטרתה לשפר את תפוקת המודל ואת רמת המגוונות שלו. הם מציגים תוצאות מעניינות של שימוש בפרקטיקה זו, ומציעים תיאור פרטי של הפרקטיקה ושל התוצאות.
תקציר מקורי באנגליתarXiv:2610.10536v1 Announce Type: new Abstract: Modern language models undergo reinforcement learning with verifiable rewards (RLVR) on top of already-trained checkpoints. A key promise of RLVR is the discovery of new reasoning strategies. In principle, a model can sample novel ideas absent from its prior training data. In practice, however, augmenting RLVR with strong novelty incentives has seen limited success and can degrade model quality. Because verifiable rewards supervise only a narrow slice of the model's knowledge and behavior, such degradations are difficult to recover from. Instead, we decouple exploration from optimization in a framework we call Exploration-Distillation (ExpDis). We train one or more explorer policies with a novelty bonus in the reward, filter their trajectorie
קרא במקור המקורי