יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

כמה יכולים מודלי שפה להרוויח מחישוב בזמן בדיקה?

How Much Can Language Models Gain from Test-Time Computation?
חוקרים בדקו כמה מודלי שפה יכולים להרוויח מחישוב בזמן בדיקה. הם השתמשו במסגרת SELF-POT כדי לבדוק את היכולת של מודלים להשתפר באמצעות חישוב נוסף. התוצאות הראו שהרווחים תלויים בתחום, כללי הבחירה וטיפול בשגיאות.
תקציר מקורי באנגליתarXiv:2610.01110v1 Announce Type: new Abstract: How much can test-time computation improve a language model, and at what cost? Test-time scaling is widely proposed as a substitute for larger models, but existing comparisons mostly evaluate one domain at a time and rarely charge selection to the budget. We introduce SELF-POT, a benchmark and evaluation framework that measures the test-time potential of a model across competition mathematics, competitive programming, and agentic workflows. SELF-POT separates candidate coverage from final accuracy on static tasks, tracks correctness transitions under revision, and measures protocol completion alongside task success in agentic environments. Under a unified budget rule, it compares Direct inference with parallel sampling and self-revision under
קרא במקור המקורי