כתבה
arXiv cs.LG ·
לא לאשלם את המודל הגדול: כיצד התפתחות סוכני תגבור משפיעה על יכולת הקוד
Don't Blame the Large Language Model: How Agent Harness Evolution Shapes Coding Agent Quality
תפתחות סוכני תגבור משפיעה על יכולת הקוד עם הזמן. חוקרים חקרו 35 גרסאות רצפיות של סוכן Qwen Code CLI ומצאו קשר ברור בין תפתחות סוכן לבין יכולת הקוד. הם גם חקרו חמש סוכני תגבור פתוחי קוד ומצאו שהם פורצלים גרסאות חדשות בקצב של יומיים.
תקציר מקורי באנגליתarXiv:2607.03691v2 Announce Type: replace-cross Abstract: Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agent harness: a middleware layer in between a developer and a large language model that orchestrates system prompts, tool execution, context management, and iterative reasoning loops. While these agent harnesses evolve at extreme velocities, no study has examined how this evolution affects agent quality (i.e., effectiveness and efficiency) over time. Practitioners regularly report quality regressions after agent harness updates, yet consistently attribute them to the underlying model rather than the harness itself. In this paper, we address this gap by conducting the first controlled longitudinal study that isolate
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית