כתבה
arXiv cs.AI ·
הערכת פלט מדויק וניבוי מצב נקודת ביקורת
Evaluating Exact Output and Checkpoint-State Prediction in Real Programs
נוצרה בנך לניבוי פלט סופי ומצב נקודת ביקורת מקוד וקלט בלבד. הבנך מכיל 400 מקרים מ-371 תוכניות Python ו-C++
תקציר מקורי באנגליתarXiv:2610.11889v1 Announce Type: cross Abstract: We present a benchmark for predicting final output and checkpoint state from source and input alone. It extends CRUXEval-style output prediction with paired shorter- and longer-trace inputs and checkpoints inside and after a loop. The benchmark contains 400 cases from 371 Python and C++ programs, evaluated under seven settings from four model families without tools or code execution. Of 11,200 planned predictions, 11,151 produced gradable responses. Reasoning-enabled settings outperform their off counterparts by 33.1 to 55.2 percentage points on completed responses. The strongest setting scores 93.0% on shorter-trace final output, 77.0% on longer-trace final output, and 65.5% and 63.5% on the two state tasks; these scores also hold when mis
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית