כתבה
arXiv cs.CL ·
Cheap to Draw, Expensive to Trust: סרטיפיקציה של קרבות סיקלינג בזמן המבחן
Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves
סרטיפיקציה של קרבות סיקלינג בזמן המבחן יקרה להיות אמינה. המאמר עוסק בפיתוח תיאור נמוך-פרמטרי של סיקלינג בזמן המבחן, ובפיתוח תיאור נמוך-פרמטרי של סיקלינג בזמן המבחן.
תקציר מקורי באנגליתarXiv:2609.40190v1 Announce Type: cross Abstract: Sampling several answers and keeping the one a verifier scores highest is one of the simplest ways to buy accuracy at test time. Its effect is reported as a scaling curve: accuracy against the number $k$ of sampled answers. The curve is cheap to draw and expensive to trust. A budget read off it is chosen after looking at every point, so only a band that covers all budgets at once protects the choice, and on a 100-question benchmark a fixed exact-binomial design needs 192,000 generated answers to certify 64 budgets to within $\pm1/32$ at 95%. Most of that cost pays for the wrong uncertainty. A benchmark is a fixed list of questions; at budget 64, about three quarters of the variance of a selected answer's correctness lies between questions,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית