כתבה
arXiv cs.LG ·
Which Self-Improvements Should We Trust? Reliable Self-Improvement When Agents Reuse Their Benchmarks
תקציר מקורי באנגליתarXiv:2609.33180v2 Announce Type: replace-cross Abstract: As recursive self-improvement (RSI) rapidly advances, reliable evaluation becomes critical for guiding adaptive search. RSI typically relies on finite evaluation resources, such as fixed benchmarks, to determine which modifications are retained and what is proposed next. However, when these finite resources are repeatedly reused, new candidates are proposed based on feedback from the same evaluation set, so the search trajectory can adaptively overfit and empirical improvement may not reflect genuine population improvement on the underlying task distribution. Some existing methods account for multiple comparisons but assume that candidates are chosen independently of the evaluation set, and therefore do not control this adaptive dep
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית