כתבה
arXiv cs.LG ·
Does Scaling Reinforcement Learning Really Require More Training?
תקציר מקורי באנגליתarXiv:2610.01133v2 Announce Type: replace Abstract: Scaling reasoning typically spends more compute on reinforcement learning (RL) or on inference. We show that a completed RL training history can yield policies stronger than the checkpoints visited by its optimizer. We call this policy-space scaling: expanding the deployable policy set accessible from a fixed RL history, without extending training or increasing per-query inference computation. We instantiate it with SURGE (Scaling Up RL Gradient-free via Eigenspace fusion). SURGE combines two checkpoints from the same RL run: a high-accuracy anchor and a competitive donor that generates shorter responses. It expresses both checkpoints as changes from their shared initialization, then spectrally decomposes the anchor's update to retain its
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית