כתבה
arXiv cs.AI ·
כשלי בטיחות תחביריים בהתפתחות הכלאה: הכרה ומעקב זמני
Compositional Safety Failures in Harness Evolution: Identification and Runtime Monitoring
חוקרים זיהו 61 כשלי בטיחות תחביריים בהתפתחות הכלאה של כליה, כולל 43 כשלי זוגיים ו-18 כשלי 3-מופעיים. הם פיתחו טכניקה של מעקב זמני שמצליחה למנוע כשלי בטיחות תחביריים.
תקציר מקורי באנגליתarXiv:2609.33123v2 Announce Type: replace Abstract: Self-evolving agent harnesses continually update persistent components such as memory, prompts, skills, and tools. We call this process harness evolution. However, such evolution could introduce unexpected safety risks. Existing work studies harness misevolution and validates candidate harnesses or attributed individual component updates, leaving safety analysis of cross-component update interactions largely unexamined. To address this gap, we study compositional safety failures in harness evolution, where interactions among individually safe and utility-preserving component updates can produce undesirable or unsafe agent behavior, revealing a safety risk intrinsic to harness evolution. Across three safety-related benchmarks, we identify
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית