כתבה
arXiv cs.AI ·
SEABench: ניטענת תקן לבדיקת תכנות עצמי באג'נטים
SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents
SEABench היא תקן לבדיקת תכנות עצמי באג'נטים, שמטרתה לחקור את הסיכון שבשיפור עצמי של אג'נטים. התקן כולל 48 רצפי משימות אורכונית, שמספרים על פני מספר פני שיפור, תחומי משימה וסוגי נזק. התקן נועד לחקור את הסיכון שבשיפור עצמי של אג'נטים, ולפתח רפרושים למניעת תכנות עצמי.
תקציר מקורי באנגליתarXiv:2609.35596v2 Announce Type: replace-cross Abstract: Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates may persist into later tasks where they produce unsafe behavior, even without direct adversarial influence. To study this risk, we introduce SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment. To account for the stochasticity inherent in agentic operati
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית