כתבה
arXiv cs.LG ·
מי מאמת את המאמן? פיתוח גרדרים נראים-לעין עם סוכנים המשכימים
Who Verifies the Verifier? Co-Evolving Inspectable Graders with Self-Improving Agents
אנו פיתחנו גרדרים נראים-לעין שמשכימים עם סוכנים, ומצאנו שהם יעילים יותר בפיתוח כישורים. המחקר נעשה על ידי פיתוח גרדרים נראים-לעין שמשכימים עם סוכנים, ומצאנו שהם יעילים יותר בפיתוח כישורים. המחקר נעשה על ידי פיתוח גרדרים נראים-לעין שמשכימים עם סוכנים, ומצאנו שהם יעילים יותר בפיתוח כישורים.
תקציר מקורי באנגליתarXiv:2610.11464v1 Announce Type: cross Abstract: We changed the agent: did it actually get better? Every self-improving agent loop answers this hundreds of times, and every answer comes from a verifier. On open-ended tasks none exists, so the loop is handed a hand-written rubric or a bare LLM judge grading output from a model like itself, inviting reward hacking and shared blind spots. We make the verifier the evolving object: an inspectable expression over small, mostly deterministic drawback detectors, synthesized from clustered failures, gated at birth, and selected for agreement with a ten-item anchored reference set plus consensus over unlabeled outputs, never for the agent's score. On MBPP+ it gains +0.21 held-out agreement over the hand-authored seed composition, on every seed, and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית