כתבה
arXiv cs.AI ·
Positive-Unlabeled Learning for Agent Safety False Alarm Auditing
תקציר מקורי באנגליתarXiv:2610.02925v1 Announce Type: new Abstract: Safety monitors help safeguard language-model agents interacting with external tools and environments, but conservative monitoring can generate many false alarms, consuming extensive review resources and weakening trust in alerts. Because false and genuine alarms often remain interleaved in native monitor scores, obtaining a reliable cutoff still requires substantial manual verification. In practice, a small set of verified-safe non-alarmed trajectories may be available while alarms remain unlabeled, naturally casting false-alarm auditing as a positive-unlabeled (PU) ranking problem. The key challenge is monitor-induced selection, since observed safe references are accepted by the monitor, while the hidden safe alarms of interest are precisel
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית