כתבה
arXiv cs.LG ·
Towards Better Exploration in Sequential Test-Time Scaling
תקציר מקורי באנגליתarXiv:2609.39632v1 Announce Type: new Abstract: Test-time scaling improves language model reasoning by spending additional compute at inference. However, both classes of existing methods often fail to continue improving over long timescales. Parallel methods repeatedly sample independent answers from the model, scaling poorly on problems the model is unlikely to solve in a single attempt. In contrast, sequential methods build on previous answers to access new ideas, yet so far have not been shown to reach answers beyond those found by parallel scaling. First, we show that sequential scaling often stops improving because it becomes prematurely trapped in an attractor: a set of answers that prevents exploration of different answers once entered. Across 27 combinations of scaling methods, mod
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית