כתבה
arXiv cs.AI ·
תיקון קוד אגנטי: חוזים לשיפור אמינות
Looping Is Not Reliability: State-Bound Evidence and Typed Revision Contracts for Agentic Code Repair
חוקרים בדקו את הפער בין תיקון קוד לאימות ואישור. הם הראו כי חזרה על תהליך התיקון לא מבטיחה אמינות. הם פיתחו חוזה טיפוסי לתיקון קוד אגנטי.
תקציר מקורי באנגליתarXiv:2607.24604v1 Announce Type: cross Abstract: Generate--test--revise loops are common in coding agents, but repetition alone provides no reliability guarantee. We study the gap between finding a correct patch and retaining, verifying, and submitting it. A sealed five-seed study over 30 HumanEval repairs produces 900 three-revision trajectories. Under forced revision, current correctness with current traces falls from 0.820 after one revision to 0.673 after two, although ever-correct rises to 0.847. Two common-state studies use 2,430 branches from identical frozen programs to remove post-treatment risk-set bias. In a prespecified 14B replication, stale traces harm 34/135 correct starts versus 4/135 with current traces, a 22.2-point increase (task-cluster 95\% CI $[8.9,37.0]$, exact Holm
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית