כתבה
arXiv cs.AI ·
Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation
תקציר מקורי באנגליתarXiv:2610.11766v1 Announce Type: cross Abstract: Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR). Yet the same non-harmful outcome can arise for very different reasons. A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely. Distinguishing these cases becomes especially important under intent-obscuring prompts, where a low ASR does not reveal whether the evaluated task was actually engaged. To make this distinction explicit, we pair ASR with operative understanding rate (UR), which measures whether a response both identifies the evaluated task and treats it as the task to be answered. Across interfaces, this paired view reveals substantial variation hidden by AS
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית