יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

טיהור של דלתא-אחורי ל-LLMs שנערכו באמצעות LoRA

Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection
החידוש: פיתוח שיטה לטיהור LLMs שנערכו באמצעות LoRA, ללא ידיעת מראש של דלתא-אחורי או עדכון מחדש. השיטה עובדת באופן יעיל ובאופן שמור, ומצליחה לצמצם את קצב ההצלחה של התקפות דלתא-אחורי לפחות מ-10%. השיטה נבדקה במספר תחומים, כולל טקסט קלאסיפיקציה והפקת תוכן.
תקציר מקורי באנגליתarXiv:2610.00685v1 Announce Type: new Abstract: With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model's general
קרא במקור המקורי