כתבה
arXiv cs.AI ·
SafeCoEvo: התפתחות בטיחות משותפת
SafeCoEvo: Co-Evolving Safety Harnesses and Guards for LLM Agents at Test-Time
SafeCoEvo היא שיטה לשיפור בטיחות של סוכנים LLM. היא מאפשרת התאמה עצמית של מערכות בטיחות בזמן ריצה. SafeCoEvo משלבת התאמה מהירה והתאמה ארוכת טווח כדי לשפר את יכולות הבטיחות.
תקציר מקורי באנגליתarXiv:2609.36580v1 Announce Type: new Abstract: LLM agents deployed in real-world environments continually encounter new tasks and safety risks, while execution feedback typically becomes available only after each task is completed. However, existing self-evolving approaches commonly rely on multiple rounds of optimization over fixed and repeatedly accessible task distributions, fundamentally differing from test-time adaptation in real-world deployment, where only experience accumulated from past tasks can be used to improve safety decisions on future unseen tasks. To address this limitation, we propose SafeCoEvo, a test-time Harness-Guard co-evolution framework for LLM agent safety that enables the external safety system to continually adapt from accumulated runtime experience. SafeCoEvo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית