כתבה
arXiv cs.AI ·
גבולות שווא: אבחון ומיתון Co-Cheating בסוכנים עצמאים
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
חוקרים גילו כי סוכנים עצמאים יכולים ליצור 'שווא' על ידי הסכמה על שגיאות משותפות. הם מציעים שיטות למיתון תופעה זו, כולל 'אימות רב-דגימות' ו-'CrossFit'.
תקציר מקורי באנגליתarXiv:2609.39102v1 Announce Type: cross Abstract: Self-evolving search agents build their own training curricula by jointly optimizing a proposer that generates questions and a solver that answers them. This closed loop introduces a failure mode we call co-cheating: the proposer and solver increasingly agree on shared errors, so internal reward improves without a matching gain in external correctness. A post-hoc audit against source evidence shows co-cheating growing more severe over successive rounds of self-evolution, with pseudo-label correctness stagnating or declining even as the in-loop training signal improves. The most direct mitigation is to verify proposals before training: we introduce multi-sample verification (MSV), which queries the same model three times with the source and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית