כתבה
arXiv cs.CL ·
RuBench: A Repository-Level Agentic Coding Benchmark with Natively Authored Russian Task Specifications
תקציר מקורי באנגליתarXiv:2607.06411v2 Announce Type: replace-cross Abstract: Developers increasingly delegate real maintenance work to product-grade coding agents, and many state tasks in their native language, in the style of a customer request rather than a curated English issue. We introduce RuBench 1.0, a benchmark of 25 tasks mined from recent fix commits in five live open-source repositories (aiohttp, aiogram, Laravel, NestJS, Fastify), each specified natively in Russian -- written from scratch, not translated -- and judged by the upstream maintainer's regression tests, which we withhold from release. All fix commits postdate the training-data cutoffs of every evaluated model. Round 1 evaluates Claude Code with Opus 4.8, Sonnet 5, and Haiku 4.5, and Codex CLI with GPT-5.5 (3 independent runs each; pass
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית