כתבה
arXiv cs.AI ·
SAGE: שער קבלה סטטיסטי לאג'נטים המתפתחים בעצמם
SAGE: A Statistical Acceptance Gate for Self-Evolving Agents
SAGE: שער קבלה סטטיסטי לאג'נטים המתפתחים בעצמם. המאמר מציג שיטה למניעת רגרסיות קבועות והטיה של האופטימיזציה. SAGE משתמש בבדיקה צולבת סטטיסטית כדי להחליט האם לקבל או לדחות שינוי. המאמר מדגים את יעילות SAGE ב-5 מבחנים שונים.
תקציר מקורי באנגליתarXiv:2609.36043v1 Announce Type: new Abstract: Large Language Model (LLM)-based agents increasingly self-evolve by editing a persistent skill document that encodes their workflow, tool-use rules, and decision logic. This loop has two steps, an optimizer that proposes a candidate edit and a gate that accepts or rejects it. Prior work has concentrated on the optimizer, while the gate still follows a naive rule that keeps any edit which improves an aggregate validation score. We show that this rule fails in two ways. First, it admits permanent regressions, since an edit can raise the average while breaking items the skill already solves. Second, it is vulnerable to the Optimizer's Curse, since the best observed score on a finite and noisy validation set is upward biased. To solve the above t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית