יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

טיהור של דלתא חשאית ב-LLMs שנערכו באמצעות LoRA

Backdoor Purification for LoRA-Tuned LLMs via Null-Space Projection
במאמר זה, המחברים מציגים שיטה לטיהור של דלתא חשאית ב-LLMs שנערכו באמצעות LoRA, כדי למנוע התקפות של דלתא חשאית.
תקציר מקורי באנגליתarXiv:2610.00685v1 Announce Type: cross Abstract: With the rapid adoption of large language models (LLMs) and parameter-efficient fine-tuning (PEFT) methods, the risk of backdoor attacks has become more severe. Existing backdoor purification methods typically rely on at least one of the strong assumptions, such as prior knowledge of triggers, access to clean references, or aggressive retraining, and they often lack comprehensive evaluations. These constraints substantially limit their practical applicability. To overcome these challenges, our work proposes purifying LoRA-tuned LLMs without these assumptions and even without post-hoc retraining of the suspect parameters. Our objective is to significantly reduce the attack success rates (ASR) while preserving both (i) the base model's genera
קרא במקור המקורי