כתבה
arXiv cs.AI ·
How to Backdoor Image Knowledge Distillation
תקציר מקורי באנגליתarXiv:2504.21323v3 Announce Type: replace-cross Abstract: Knowledge distillation is widely used to transfer behavior from a large teacher model to a smaller student. It is often assumed to be safe when the teacher is clean, because classic backdoor attacks rely on poisoned labels and triggers in supervised training, whereas distillation trains the student to match a teacher's outputs. We show that this assumption can fail when the distillation dataset itself is poisoned. Our attack injects triggered and manipulated images that a clean teacher already predicts as an attacker chosen target label, which causes the student to learn a backdoor even though the teacher remains unaffected. We evaluate this threat across multiple manipulation strategies, including targeted adversarial perturbations
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית