יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

TestJack: בדיקת תוצאות בבנכמרקרים של קידוד

TestJack: Should you trust the results in coding benchmarks? Agentic Coding Benchmarks Auditing via Evaluator Evolution
TestJack הוא כלי לבדיקת פתרונות קידוד. הוא מזהה פתרונות שעוברים מבחנים אך אינם עומדים בדרישות המשימה. TestJack נוסה על 6 מודלים ו-5 בנכמרקרים, והראה ש-34.4% מהפתרונות הנחשבים נכונים אינם עומדים בדרישות.
תקציר מקורי באנגליתarXiv:2610.10619v1 Announce Type: cross Abstract: Large language model (LLM) agents are rapidly reshaping software engineering, accompanied by an explosion of new code benchmarks. Yet nearly all existing benchmarks still rely on the same decades-old criterion: a solution is correct if it passes a fixed set of unit tests. Such tests are often insufficient: they check only part of what the task requires, so agents can reward hack them or silently miss required behavior while still passing every test. As a result, higher benchmark scores may partly reflect better adaptation to the evaluator rather than better problem solving. Existing works focus on static test augmentation: they strengthen each task's tests once, before any trial is seen, and thus overlook how real trials actually fail. We i
קרא במקור המקורי