יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

The Model Plants the Trigger: Answer-Side Backdoor Attacks in Multi-Turn Large Language Models

תקציר מקורי באנגליתarXiv:2610.07723v1 Announce Type: cross Abstract: Safety alignment in Large Language Models (LLMs) remains vulnerable to backdoor attacks. Existing LLM backdoors are almost all input-centric: activation depends on explicit trigger patterns in the user input, so modern guardrails are built to sanitize the input space. We challenge this assumption with a novel answer-side backdoor for multi-turn dialogue. Instead of inserting the trigger into the input, the adversary uses a benign first-turn prompt to naturally induce the model to generate a specific, seemingly innocuous word. Once merged into the dialogue history, this self-generated word becomes the trigger. When a later harmful query arrives, the model detects its own trigger and bypasses its safety refusal, while the user input stays per
קרא במקור המקורי