כתבה
arXiv cs.AI ·
Groundability, Not Scale Alone: When Weak Reviewers Can Audit Strong Coding Agents
תקציר מקורי באנגליתarXiv:2610.01023v1 Announce Type: cross Abstract: Coding agents can return plausible patches that omit required behavior. These failures are hard to review because long traces and confident summaries often hide what was missed. We ask when a nominally weaker reviewer can reliably decide whether a patch solves its issue. We study 411 execution-labeled traces from three agents and 101 controlled cases. On 154 GPT-5.4 traces, structured but unchecked evidence raises both defect catch and over-rejection. We then provide official execution evidence as an upper-bound diagnostic. After choosing and freezing one of two formats per reviewer, five of six reviewers improve both rates on 122 held-out traces; two classify every trace correctly. Reviewer size is not a consistent predictor of quality. Be
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית