כתבה
arXiv cs.AI ·
האם גידול בלמידת רפלקסיה דורש יותר הכשרה?
Does Scaling Reinforcement Learning Really Require More Training?
לא נדרשת הכשרה נוספת כדי לגדול את למידת הרפלקסיה. היסטוריה שלמה של RL יכולה לתת פוליציות חזקות יותר. ניתן לשלב שני נקודות ציון מההיסטוריה כדי ליצור פוליציות חזקות יותר. ניתן לבחור את גודל הבלוק על פי המשקל של המודל. ניתן לבחור את הפוליציות שיש להציג.
תקציר מקורי באנגליתarXiv:2610.01133v1 Announce Type: cross Abstract: Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to retain its d
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית