כתבה
arXiv cs.AI ·
Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
תקציר מקורי באנגליתarXiv:2610.11464v1 Announce Type: new Abstract: We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and shared blind spots. We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. On MBPP+ it gains +0.21 held-out agreement over the hand-authored seed composition, on every seed, and e
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית