כתבה
arXiv cs.AI ·
האם גידור של RL דורש יותר הכשרה?
Does Scaling Reinforcement Learning Really Require More Training?
אנו מראים שהיסטוריה של RL שלמה יכולה לתת מדיניות חזקה יותר מהנקודות-מפקד של האופטימיזציה. אנו מציגים את SURGE, שמשלב שני נקודות-מפקד מאותה הכשרה: אחד עם דיוק גבוה ואחד עם תגובות קצרות. SURGE משמר את המרכיב הדומיננטי של העדכון של הנקודת-המפקד הגבוהה ומשלב את המרכיב המשלים של הנקודת-המפקד התחרותי. SURGE משפר את דיוק המדיניות על שני הנקודות-מפקד המקוריות, תוך שימוש בפחות יחידות-הסברה מאשר הנקודת-המפקד הגבוהה.
תקציר מקורי באנגליתarXiv:2610.01133v2 Announce Type: replace-cross Abstract: Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to reta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית