יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בחרו את הרע: למידה של בחירת קבוצות רע להתקפי LLM חזקים

Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
במאמר זה, נראה כיצד ניתן לבחור קבוצות רע כדי ליצור התקפי LLM חזקים. המחברים פיתחו שיטה חדשה לבחירת קבוצות רע, הקרויה SAILS, שמשפרת את הצלחת ההתקפה ב-30%.
תקציר מקורי באנגליתarXiv:2609.15029v1 Announce Type: cross Abstract: Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poisoned examples and sample them at random from a candidate pool. We show that this can severely underestimate worst-case vulnerability: across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen. We formalize poison selection as oracle-budgeted set optimization and introduce SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-a
קרא במקור המקורי