כתבה
arXiv cs.AI ·
SWE-Journey: דרך יותר אותנטית לביקורת עוזרי קוד
SWE-Journey: Towards More Realistic Evaluation of Coding Assistants through Long-Horizon, Multi-Turn Interaction
במאמר זה נציגים תקן חדש לביקורת עוזרי קוד, המתמקד בשיחות ארוכות-טווח ובשיחות רב-סיבוב. התקן, הנקרא SWE-Journey, נועד לבדוק את יכולתם של עוזרי קוד לסייע למתכנתים בפיתוח תוכנה. המחברים טענו כי עוזרי קוד הנוכחיים עדיין לא יכולים לספק סיוע יעיל למתכנתים שאינם מקצועיים.
תקציר מקורי באנגליתarXiv:2610.11559v1 Announce Type: cross Abstract: Coding assistants such as Claude Code and Codex have become a major application of LLM agents, yet existing benchmarks remain far from real-world use, particularly in task horizon and interaction length. Code assistants require completing long chains of development work in continuously evolving repositories, while repeatedly clarifying requirements and adapting implementations through multi-turn interaction. To address these gaps, we introduce SWE-Journey, a benchmark for more realistic evaluation of coding assistants. To address the task-horizon gap, we propose a weak-to-strong synthesis pipeline that automatically constructs long-horizon coding tasks. To address the interaction gap, we mine four representative user personas from real inte
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית