כתבה
arXiv cs.LG ·
Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves
תקציר מקורי באנגליתarXiv:2609.40190v1 Announce Type: new Abstract: Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is reported as a scaling curve: accuracy against the number $k$ of sampled answers. The curve is cheap to draw and expensive to trust. A budget read off it is chosen after looking at every point, so only a band that covers all budgets at once protects the choice, and on a 100-question benchmark a fixed exact-binomial design needs 192,000 generated answers to certify 64 budgets to within $\pm1/32$ at 95%. Most of that cost pays for the wrong uncertainty. A benchmark is a fixed list of questions; at budget 64, about three quarters of the variance of a selected answer's correctness lies between questions, an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית