כתבה
arXiv cs.LG ·
הכלאה של נתיבי נגישה ב-LLMs
Backdoor Containment via Expert Quarantine and Shutdown in LLMs
אנשי מחקר הציגו שיטה חדשה למניעת נתיבי נגישה ב-LLMs. השיטה, הנקראת QES, מאפשרת למודלים להיות פתוחים לאימון על נתיבי נגישה, אך לא להפעיל אותם בשלב השימוש. QES נבנה על בסיס מודלי Transformer ומשתמש בטכניקות של LoRA ורוטינג.
תקציר מקורי באנגליתarXiv:2610.00663v1 Announce Type: cross Abstract: Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, learn, but channel: allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose Quarantined Expert Shutdown QES,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית