כתבה
arXiv cs.AI ·
Active-SWE: Benchmarking Coding Agents for Proactive Bug Fixing without Issue Reports
תקציר מקורי באנגליתarXiv:2608.04682v2 Announce Type: replace-cross Abstract: Coding agents powered by large language models (LLMs) are increasingly adopted in software engineering (SWE) scenarios, capable of fixing a specific bug in large-scale codebase. However, existing SWE benchmarks typically assume that high-quality issue reports with detailed information are always available, which is easily violated in practice due to the complexity of report acquisition and curation. To address this, we introduce Active-SWE, a benchmark for evaluating coding agents on proactively discovering and fixing multiple bugs without report guidance, covering 1,663 tasks across six bug categories and eight languages. Beyond shifting the focus from existing reactive bug fixing to proactive bug fixing, Active-SWE enables a more
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית