כתבה
arXiv cs.AI ·
Evaluating System One Models for Agent Security Decisions: Reliability, Calibration, and Selective Automation
תקציר מקורי באנגליתarXiv:2609.33401v2 Announce Type: replace-cross Abstract: Model-based judges support agent security by detecting prompt injections, assessing interaction risks, and screening harmful requests. System One models select from predefined answers and report probabilities that software can use to allow, block, or review inputs, but the reliability of these automated decisions remains unclear. We evaluate Jev, Laya, Decider, and Bespoke Nimble against specialized classifiers and language-model judges, examining decision accuracy, probability calibration, and selective automation. We draw the following conclusions. (1) Strong overall performance and favorable aggregate calibration can hide failures concentrated in particular attack groups, including attacks classified as safe with high confidence.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית