יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

MTAC-IFBench: בנצ'מרק למעקב אחר הוראות בקידוד אגנטי רב-תור

MTAC-IFBench: Benchmarking Instruction-Following in Multi-Turn Agentic Coding
MTAC-IFBench הוא בנצ'מרק לבדיקת יכולת המעקב אחר הוראות בקידוד אגנטי רב-תור. הוא כולל הוראות פיתוח תוכנה רב-תוריות עם אילוצים מגוונים. הבנצ'מרק מציג אתגר קשה למודלים הקיימים, ומזהה חוסרים משמעותיים ביכולת המעקב אחר הוראות בקידוד אגנטי רב-תור.
תקציר מקורי באנגליתarXiv:2609.14992v1 Announce Type: new Abstract: Recently, the rapid development of large language models (LLMs) has reshaped software engineering by enabling autonomous code agents that plan, execute, and utilize external tools iteratively to tackle complex tasks. Beyond achieving functional correctness, these agents must faithfully follow process instructions and constraints throughout the development lifecycle. However, existing benchmarks typically focus on final functional correctness or confine instruction-following evaluation to single-turn, general chat or simple code generation scenarios, leaving instruction-following in multi-turn agentic coding underexplored. To bridge this gap, we propose MTAC-IFBench, a comprehensive benchmark for this critical capability. It features multi-tur
קרא במקור המקורי