יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

יציבות סמנטית של LLM

How Semantically Stable Are LLM Refusals? Measuring Confusion in Local Safety Boundaries
חוקרים פיתחו מדדים חדשים לבדיקת יציבות סמנטית במודלים LLM. המחקר בודק כיצד מודלים כמו Llama מגיבים לבקשות דומות. התוצאות מראות שיש לשפר את היציבות הסמנטית במודלים אלו.
תקציר מקורי באנגליתarXiv:2512.01037v3 Announce Type: replace-cross Abstract: As safety alignment becomes standard in large language models, refusal behavior has become an important part of model reliability. However, models may still reject benign prompts, especially when the wording resembles risky content. Existing evaluations usually report global scores, such as false rejection rate or compliance rate. These scores are useful, but they treat each prompt independently. As a result, they miss local inconsistency, where a model accepts one phrasing of an intent but rejects a close paraphrase. This makes it difficult to understand whether refusals are only frequent or also semantically unstable. We address this gap by introducing Semantic Confusion, a failure mode that captures contradictory refusal decision
קרא במקור המקורי