כתבה
arXiv cs.CL ·
מדידה ומיתון קריסת מצב פתרון ב-RLVR
Measuring and Mitigating Solution Mode Collapse in RLVR
RLVR הוא שיטה ללמידת חיזוק עם פרסים מאומתים. מחקר זה מציג את ModeBench, בנק אימות למשימות רב-פתרון, ואת Re:Max, שיטה לשימור גיוון פתרונות. התוצאות מראות ש-Re:Max משפרת את הצלחת המדיניות ואת מספר הדרכים להצליח.
תקציר מקורי באנגליתarXiv:2610.11064v1 Announce Type: cross Abstract: A language model (LM) can usually answer the same question in more than one way, but reinforcement learning with verifiable rewards (RLVR) is indifferent to which correct answer a model produces. A solution will earn the same reward whether it is the thousandth copy of a familiar answer or one the model has never produced before. Yet, there is potential value in having the model retain multiple correct solutions as it is trained. For instance, multiple modes may give users a choice and provide problem-solving strategies that improve overall model performance. Here, we introduce ModeBench, a benchmark of multi-solution tasks in which the verifier returns both correctness and mode discovered. We then use ModeBench to measure how solution dive
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית