כתבה
arXiv cs.LG ·
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
תקציר מקורי באנגליתarXiv:2609.39102v1 Announce Type: cross Abstract: Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית