יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SafeCoEvo: פיתוח בזמן הריצה של חימרים לבטיחות LLM

SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time
SafeCoEvo היא תשתית לבטיחות LLM שמאפשרת למערכת הבטיחות להתאים בזמן הריצה למשימות חדשות. היא משלבת שני יכולות בטיחותיות: S-Harness, שמתעדכן בזמן הריצה, ו-GuardVPO, שמתעדכן בזמן הריצה ומשפר את יכולות הבטיחות של האגן.
תקציר מקורי באנגליתarXiv:2609.36580v2 Announce Type: replace Abstract: LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCo
קרא במקור המקורי