יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

תוצאה זהה, ראיות שונות: השבתת מטרה בבדיקת בטיחות LLM

Same Outcome, Different Evidence: Intent Recovery in LLM Safety Evaluation
בדיקות בטיחות של מודלי שפה גדולים משתמשות לעיתים קרובות בקצב הצלחה של התקפה (ASR).
תקציר מקורי באנגליתarXiv:2610.11766v1 Announce Type: new Abstract: Safety evaluations of large language models commonly summarize harmful-output behavior with attack success rate (ASR). Yet the same non-harmful outcome can arise for very different reasons. A model may recover a harmful task and refuse it, fail to recover the task, or respond to something else entirely. Distinguishing these cases becomes especially important under intent-obscuring prompts, where a low ASR does not reveal whether the evaluated task was actually engaged. To make this distinction explicit, we pair ASR with operative understanding rate (UR), which measures whether a response both identifies the evaluated task and treats it as the task to be answered. Across interfaces, this paired view reveals substantial variation hidden by ASR:
קרא במקור המקורי