כתבה
arXiv cs.AI ·
בחירת הרעל: לימוד לבחור קבוצות רעל לתקיפות Backdoor חזקות יותר
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
SAILS היא שיטה חדשה לבחירת קבוצות רעל לתקיפות Backdoor על מודלים כמו LLaMA. השיטה משפרת את הצלחת התקיפות על ידי לימוד דירוג קבוצות רעל ובחירת הטובות ביותר. המחקר מראה כי SAILS משפרת את הצלחת התקיפות ב-30% לעומת שיטות אחרות.
תקציר מקורי באנגליתarXiv:2609.15029v1 Announce Type: cross Abstract: Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poisoned examples and sample them at random from a candidate pool. We show that this can severely underestimate worst-case vulnerability: across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen. We formalize poison selection as oracle-budgeted set optimization and introduce SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית