כתבה
arXiv cs.LG ·
How Much Can Language Models Gain from Test-Time Computation?
תקציר מקורי באנגליתarXiv:2610.01110v2 Announce Type: replace Abstract: How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that measures the test-time potential of a model across competition mathematics, competitive programming, and agentic workflows. SELF-POT separates candidate coverage from final accuracy on static tasks, tracks correctness transitions under revision, and measures protocol completion alongside task success in agentic environments. Under a unified budget rule, it compares Direct inference with parallel sampling and self-revision u
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית