יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

בנצ'מרקים לסוכני קידוד: התאמה לזרימת משימות

Coding-Agent Benchmarks Should Match Their Users' Task Flows
חוקרים פרסמו מחקר על בנצ'מרקים לסוכני קידוד. הם גילו שהערכה ריאליסטית דורשת התאמה לזרימת משימות של משתמשים. הם הציגו גישה חדשה, SWE-TaskFlow, לשיפור הבנצ'מרקים.
תקציר מקורי באנגליתarXiv:2610.09633v1 Announce Type: new Abstract: The evaluation of coding agents generally strives to be as realistic as possible. In our study, we collect 4,782 agent sessions of real software engineers in JetBrains IDEs, which we call Production Sessions. Since our subject is interactive agents, we study the sessions with at least three user messages (33% of the sample). These long sessions differ from issue-derived benchmark tasks in two ways: (i) user requests span a far wider mix of task types - questions about the project's code, planning, review, refactoring, execution - and (ii) users switch between types throughout a session. Long-session samples from three public interaction corpora exhibit markedly different Task Flows (the distributions of session lengths, task types, and type-t
קרא במקור המקורי