כתבה
arXiv cs.AI ·
E2E-SWE: בדיקת מודלי LLM לבניית קוד תפעולי מאפס
E2E-SWE: Benchmarking LLMs on Building Working Codebases from Scratch
בדיקת מודלי LLM לבניית קוד תפעולי מאפס. המחקר מציג בסיס נתונים חדש (E2E-SWE) לבדיקת יכולות של מודלי LLM לבניית קוד תפעולי מאפס. הבדיקה כוללת 186 תרגילים של בניית קוד תפעולי מאפס ב-11 שפות תכנות. המחקר מציג תוצאות של 13 מודלי LLM, ומציע רעיונות לשיפור יכולות המודלים.
תקציר מקורי באנגליתarXiv:2609.38335v1 Announce Type: cross Abstract: Coding agents powered by large language models (LLMs) are evolving from making localized code changes to developing complete software repositories. However, evaluating repository-scale generation remains challenging: tasks must demand system-level reasoning while ensuring that all evaluated behaviors are precisely specified and independent of any particular implementation. We introduce E2E-SWE, a benchmark for evaluating whether coding agents can build complete, functional software repositories end to end. E2E-SWE contains 186 whole-repository generation tasks spanning 11 programming languages. Given only a natural-language specification and an empty workspace, an agent must implement a complete, installable project that satisfies a compreh
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית