כתבה
arXiv cs.LG ·
על הנרתיקות של סוכני השיפור העצמי: תחושתיות, סדר תפקידים ולא-מפורשות
On the Fragility of Self-Improving Agents: Variance, Task Order, and Underspecification
נמצאו חולשות בסוכני השיפור העצמי, שהתגלו בניסויי תחושתיות וסדר תפקידים. החוקרים חקרו את השפעת סדר תפקידים על סוכני השיפור העצמי ומצאו כי השיפור של הסוכן תלוי בסדר תפקידים. החוקרים גם חקרו את השפעת תחושתיות על סוכני השיפור העצמי ומצאו כי השיפור של הסוכן תלוי בתחושתיות. החוקרים חקרו גם את השפעת לא-מפורשות על סוכני השיפור העצמי ומצאו כי השיפור של הסוכן תלוי בלא-מפורשות.
תקציר מקורי באנגליתarXiv:2608.18066v2 Announce Type: replace-cross Abstract: Memory-based self-improving agents--those that learn from an online stream of tasks and improve over time by maintaining a textual memory bank--have shown great promise in recent literature. However, the reliability aspects of these methods have been critically overlooked. In this work, we conduct a comprehensive re-evaluation of two memory-based methods, broadening the scope of evaluation along two axes: (1) including multiple self-improving runs to quantify variance, and (2) shuffling the tasks to investigate the effect of task order. Through these experiments, we make two observations that expose the fragility of current methods: First, agent evaluation is inherently noisy in complex environments and on multi-step tasks, and stac
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית