כתבה
arXiv cs.AI ·
הכלאה של Backdoor ב-LLM
Backdoor Containment via Expert Quarantine and Shutdown in LLMs
חוקרים הציעו שיטה חדשה להכלאת Backdoor במודלים גדולים של שפה. השיטה, Quarantined Expert Shutdown, מאפשרת להכליא את ה-Backdoor ברכיב מיוחד שניתן לנטרל בעת הפריסה. השיטה הוכחה כיעילה במניעת התקפות Backdoor במודלים שונים.
תקציר מקורי באנגליתarXiv:2610.00663v1 Announce Type: new Abstract: Backdoored large language models (LLMs) can behave normally on benign inputs while producing attacker-specified outputs under hidden triggers. Existing defenses span four stages--prior-training, in-training, post-training, and inference-time--and share one of two underlying strategies: either suppress backdoor learning (by filtering poisoned data or interrupting its acquisition during optimization) or learn, then purify (by repairing model weights or gating inputs after a fully backdoored model has formed). We propose a third strategy, learn, but channel: allow backdoor formation during training but route it into a designated, quarantined component that can be disabled at deployment. To this end, we propose Quarantined Expert Shutdown QES, a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית