כתבה
arXiv cs.LG ·
Pick Your Poison: Learning to Select Poison Sets for Stronger LLM Backdoor Attacks
תקציר מקורי באנגליתarXiv:2609.15029v1 Announce Type: new Abstract: Backdoor poisoning attacks add poisoned examples to otherwise-clean finetuning data, pairing a trigger with a target behavior that the model learns to produce when the trigger appears. Existing evaluations typically fix the number of poisoned examples and sample them at random from a candidate pool. We show that this can severely underestimate worst-case vulnerability: across three LLaMA-3-8B backdoor settings, holding the model, clean data, and poison count fixed, attack success ranges from 3% to 80% depending only on which poison set is chosen. We formalize poison selection as oracle-budgeted set optimization and introduce SAILS (Set-level Audit-Informed Iterative Learned Selection), which learns a set scorer from a few hundred finetune-and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית