כתבה
arXiv cs.CL ·
Self-Authored Verification Is Unreliable in Heuristic Self-Improving Agents
תקציר מקורי באנגליתarXiv:2607.24300v1 Announce Type: new Abstract: Self-improving agents accumulate capability by repeatedly rewriting procedural policies, controllers, or heuristic rules. They typically rely on self-authored tests or metrics to decide whether to accept subsequent edits. The agent controls both the optimized object and its verifier. As a result, self-assigned scores can remain near perfect while real deployment performance degrades or stays low. We study this problem through the verifier--deployment gap. This gap refers to the discrepancy between an agent's self-authored verification signal and a sealed deployment evaluation that the agent cannot observe or access. We ask how self-authored verification fails under iterative policy-and-test rewriting, how the failure changes with capability,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית