יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

הסרת המחט מהעשב: הסרת Backdoor ב-LLMs

Removing the NEEDLE in the Haystack: Backdoor Removal in LLMs via Weight Orthogonalisation
NEEDLE היא שיטה חדשה להסרת Backdoor ממודלי LLM. השיטה מאפשרת הסרה מהירה ויעילה של Backdoor מבלי לפגוע בביצועים של המודל. NEEDLE נבדקה על מספר מודלים והשיגה תוצאות מרשימות.
תקציר מקורי באנגליתarXiv:2610.00348v1 Announce Type: cross Abstract: Backdoor attacks can be implanted in Large Language Models (LLMs) during training, causing unwanted behaviour when a trigger appears in the input. Existing backdoor defences for LLMs attempt to remove the backdoor but inadvertently shift the model's output distribution to benign prompts, which can result in degraded model performance and safety. We propose NEEDLE, a training-free method for targeted backdoor removal. Once a trigger has been identified, our method estimates a backdoor direction and a refusal subspace through activation vectors, then applies sequential weight orthogonalisation to suppress the backdoor while preventing changes in refusal-related representations. NEEDLE requires neither a clean reference model nor the original
קרא במקור המקורי