יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

לכיוון מוניטורינג יחיד של שימוש לא ראוי

Towards a Unified Misuse Monitoring Benchmark
מאמר חדש עוסק בפיתוח מדד יחיד למוניטורינג שימוש לא ראוי של כלי למידת מודלים. המדד כולל תרגילים של קטעי שיחה בין משתמש, כלי LLM ומערכת חיצים. המאמר גם מציג תוצאות של 17 תצורות של מוניטורינג שהוצגו במדד.
תקציר מקורי באנגליתarXiv:2610.07089v1 Announce Type: cross Abstract: LLM agents increasingly act in multi-actor environments, exposing them to misuse from multiple sources: decomposition attacks, where a harmful request is split into innocuous sub-requests, and prompt injection attacks, where a compromised tool delivers a malicious instruction. Existing evaluations treat these threats separately and ask whether a trajectory is harmful, rather than when it becomes harmful. We propose monitoring the agent's responses, where its actions are externalised, and ask whether the first point where monitors identify harm lands within a harm window (from the agent's first harmful commitment to goal execution). We develop a unified formalism for trace-level misuse monitoring and use it to construct a benchmark of ~6,200
קרא במקור המקורי