כתבה
arXiv cs.CL ·
Safety Monitors Mostly Catch What the Model Already Refuses
תקציר מקורי באנגליתarXiv:2609.05797v3 Announce Type: replace Abstract: Safety monitors are evaluated by recall on harmful prompts, regardless of whether the target model would answer them. Yet a monitor matters most on the prompts the model does answer. We measure recall on exactly those prompts, defined by sampling the target model and judging its responses. Across four text guards, two activation probes, and Latent Guard, recall at a 1% false positive rate falls sharply on this subset: at a common threshold, every monitor catches the requests the model refuses 1.1 to 6.4 times as often as the requests it answers. Standard metrics hide this; AUROC stays above 0.85 for most monitors. Rewriting each request to be less explicit, with intent held fixed and verified, raises compliance 28-fold and lowers every mo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית