כתבה
arXiv cs.CL ·
עד כמה יציבים LLM בסירוב? מדידת הבלבול בגבולות הבטיחות
How Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries
במאמר זה נחקרה יציבות ה-LLM בסירוב, ונצפה כיצד הבלבול נפוץ בגבולות הבטיחות. המחברים הציגו תרחישים של סירוב שגוי, והציעו תרחישים חדשים לבדיקת יציבות ה-LLM.
תקציר מקורי באנגליתarXiv:2512.01037v3 Announce Type: replace Abstract: As safety alignment becomes standard in large language models, refusal behavior has become an important part of model reliability. However, models may still reject benign prompts, especially when the wording resembles risky content. Existing evaluations usually report global scores, such as false rejection rate or compliance rate. These scores are useful, but they treat each prompt independently. As a result, they miss local inconsistency, where a model accepts one phrasing of an intent but rejects a close paraphrase. This makes it difficult to understand whether refusals are only frequent or also semantically unstable. We address this gap by introducing Semantic Confusion, a failure mode that captures contradictory refusal decisions acro
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית