כתבה
arXiv cs.CL ·
SEABench: בדיקת הסטייה העצמית בסוכנים המתפתחים
SEABench: Benchmarking Endogenous Misalignment In Self-Evolving Agents
SEABench הוא בנץ'מרק לבדיקת הסטייה העצמית העלולה להתרחש בסוכנים המתפתחים אוטונומית. המחקר בודק 48 רצפי משימות שונים, ומראה כי התפתחות עצמית משפרת את שיעורי ההצלחה, אך גם גורמת לתקלות בטיחות. הממצאים מראים כי התנהגות הבטיחות משתנה בין סוכנים שונים וסוגי נזק.
תקציר מקורי באנגליתarXiv:2609.35596v2 Announce Type: replace-cross Abstract: Self-evolving LLM agents have gained prominence for their ability to improve after deployment by modifying their harness, including their controller instructions, memory management protocols, and reusable tools and skills, in response to user and environment feedback. However, locally useful updates may persist into later tasks where they produce unsafe behavior, even without direct adversarial influence. To study this risk, we introduce SEABench, a benchmark for studying endogenous misalignment arising from agent self-evolution, with 48 longitudinal task sequences that span multiple evolution surfaces, task domains, and harm types in a rich personal-assistant environment. To account for the stochasticity inherent in agentic operati
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית