יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מיתון סירוב יתר במודלים גדולים

Mitigating LLM Over-Refusal via Dynamic Semantic Routing Calibration
חוקרים גילו כי מודלים גדולים סובלים מסירוב יתר. הם הציעו אלגוריתם חדש למיתון סירוב יתר, Semantic Routing Calibration, שמשפר את היכולת של המודלים להבין הוראות בטוחות.
תקציר מקורי באנגליתarXiv:2609.25049v2 Announce Type: replace-cross Abstract: Large language models (LLMs) aligned for safety often suffer from over-refusal, incorrectly rejecting benign yet safety-related instructions. Prior studies primarily attribute this to static representation overlap, largely overlooking the underlying dynamic mechanisms. In this paper, we present the mechanistic analysis of over-refusal through the lens of internal routing conflicts within transformer attention. We discover that a sparse subset of Hypersensitive Safety Heads misfires on Hard-Safe prompts, exhibiting abnormal attention entanglement that forcefully binds harmless target entities to refusal semantics. This triggers a severe, high-entropy routing conflict that deprives target entities of necessary attention. To counteract
קרא במקור המקורי