יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

האם הגדלת למידת חיזוק דורשת אימון נוסף?

Does Scaling Reinforcement Learning Really Require More Training?
חוקרים מציגים שיטה חדשה לשיפור ביצועי מודלים של למידת חיזוק, בלי צורך באימון נוסף. השיטה, הנקראת SURGE, משלבת שני נקודות צ'קפוינט מאותו רץ אימון, ומשפרת את הדיוק של המודל. החוקרים בדקו את השיטה על מודלים שונים, כולל DeepSeek, וקבלו תוצאות משופרות.
תקציר מקורי באנגליתarXiv:2610.01133v1 Announce Type: new Abstract: Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to retain its dom
קרא במקור המקורי