כתבה
arXiv cs.LG ·
לא תדירות, אלא רגישות: זיהוי מומחים בטוחים ב-LLM רב-מומחים דקת-זמן
Frequency Is Not Sensitivity Identifying Safety-Sensitive Experts in Sparse MoE LLM
במאמר זה, חוקרים חוקרים את האפקט של צמצום קבוצת מומחים ב-LLM רב-מומחים דקת-זמן על תפקודי בטיחות. הם מציעים תחליף למדד התדירות, ומצביעים על יתרונותיו של רגישות-גרדיאנט.
תקציר מקורי באנגליתarXiv:2610.02910v1 Announce Type: cross Abstract: Suppressing a small set of routed experts can weaken the safety behavior of a sparse Mixture-of-Experts (MoE) language model without retraining. Which experts to suppress is therefore a security question, and the usual answer is activation frequency, but frequency measures use, not influence. We test an alternative: router-gradient sensitivity, the sensitivity of the sequence loss to the gate weights that select an expert. Across five MoE architectures, we rank experts by each signal on 500 benign and 500 malicious prompts and measure refusal on 100 held-out malicious prompts under two budgets: equal expert counts and equal nominal malicious routing traffic (1%-5%). Under each of the two budgets, router-gradient selection reduces refusals m
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית