כתבה
arXiv cs.LG ·
לכיוון ניטרוליזציה של מערכת המשמרת
Towards a Unified Misuse Monitoring Benchmark
מערכת ניטרוליזציה חדשה ל-LLMs, שמטרתה לזהות פעולות מסוכנות. המערכת נבחנה ב-6,200 תקצירי שיחה בין גורם חיצוני, LLM ומשתמש.
תקציר מקורי באנגליתarXiv:2610.07089v1 Announce Type: cross Abstract: LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluations treat these threats separately and ask whether a trajectory is harmful, rather than when it becomes harmful. We propose monitoring the agent's responses, where its actions are externalised, and ask whether the first point where monitors identify harm lands within a harm window (from the agent's first harmful commitment to goal execution). We develop a unified formalism for trace-level misuse monitoring and use it to construct a benchmark of ~6,200
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית