יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

כל עבודה ואין שעשועה: הבנת ומניעת קריסת סטרטגיה קטסטרופלית ב-RLVR

All Work And No Play Makes Jack a Dull Boy: Understanding and Preventing Catastrophic Strategy Collapse in RLVR
במאמר זה, המחברים חוקרים את הקריסת סטרטגיה הקטסטרופלית ב-RLVR ומציעים פתרון למניעתה. הם מציגים תאוריה יחידה שמשלבת דינמיקות אופטימיזציה ותורת המידע. הם גם מציעים את 'Mesh Learning' כפתרון למניעת קריסת סטרטגיה.
תקציר מקורי באנגליתarXiv:2610.02835v1 Announce Type: new Abstract: During post-training of large language models (LLMs) with Reinforcement Learning with Verifiable Rewards (RLVR), GRPO-style algorithms can exhibit severe late-stage collapse. Prompt-based probing reveals that this is not benign strategic pruning, but a harmful contraction of effective strategy capacity that makes distinct reasoning strategies increasingly inaccessible. To characterize this phenomenon, we define strategies through trajectory-level policy-update interactions and develop a unified theoretical framework combining optimization dynamics and information theory. We prove that major RLVR objectives progressively concentrate probability mass onto a single strategy, while sustaining nontrivial task accuracy requires a minimum strategy c
קרא במקור המקורי