כתבה
arXiv cs.CL ·
Position: Coding Benchmarks Are Misaligned with Agentic Software Engineering
תקציר מקורי באנגליתarXiv:2606.17799v2 Announce Type: replace-cross Abstract: Coding agents have become a major mode of software engineering, but the benchmarks we use to compare them were designed in a pre-agent era: they collapse model, harness, and environment into a single end-to-end score, typically computed against one reference solution, with no component-level signal for iteration. We argue that current coding benchmarks are misaligned with agentic software engineering. A coding agent in practice is not a model: it is a system harness -- a composite of models, harnesses, contexts, environments, and feedback signals, any one of which can move the benchmark score by margins comparable to those between adjacent model generations. We discuss three symptoms: (i) benchmark scores conflate the model with the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית