כתבה
arXiv cs.CL ·
לפתוח את הבלתי ניתנים: תוכנית לימודים מונחה על ידי מורה ללמידה עם תוצאות רווח
Unlocking the Unsolvable: Teacher-Guided Curriculum for Data-Efficient RLVR
אנו מציגים תוכנית לימודים מונחה על ידי מורה שמאפשרת ללמידת מודלים לפתוח בעיות שלא ניתן היה לפתור קודם. התוכנית כוללת רצף סיבות מוחלט של מודל חזק, ומתקדם באופן התקדמותי עד שהמודל פותר בעיות בעצמו. התוצאות היו טובות יותר מאשר GRPO שהוכשר על קורפוס שלם של 2,000 בעיות.
תקציר מקורי באנגליתarXiv:2609.13997v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has shown remarkable success in improving the mathematical reasoning of large language models. Yet problems beyond the model's current capability, where rollouts uniformly fail and no learning signal is produced, are structurally wasted despite marking the most informative training frontier. We show that these otherwise-inert problems can be unlocked via teacher-guided curriculum learning: partial reasoning traces from a stronger model create a graded difficulty landscape, and a backward-chaining curriculum progressively withdraws guidance until the model solves problems unaided. Training on only 128 unsolvable problems matches or exceeds GRPO trained on a full 2,000-problem corpus (~16x d
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית