כתבה
arXiv cs.AI ·
Bait-and-Recover: פגיעה באובייקטים והשבה - הגנה נגד עיצוב ישיר
Bait-and-Recover: Poisoning Internal Refusal Signals to Defend LLMs against White-Box Editing Jailbreaks
Bait-and-Recover: פגיעה באובייקטים והשבה - הגנה נגד עיצוב ישיר של מודלי LLM. המחברים הציעו פתרון לבעיה של עיצוב ישיר של מודלי LLM. הם פיתחו טכניקה חדשה שנקראת Bait-and-Recover, שמטרתה להגן על מודלי LLM מפני עיצוב ישיר.
תקציר מקורי באנגליתarXiv:2609.05794v1 Announce Type: cross Abstract: Open-weight large language models face a low-cost white-box threat from representation engineering attacks. Attackers can estimate refusal directions and search for projection-matrix edits that suppress safety alignment while preserving general capabilities, within minutes on a single GPU and without gradient-based training. We propose Bait-and-Recover, a weight-level defense that places a bait adapter where attackers read activations and a paired recovery adapter at the subsequent layer. Trained via gradient routing, this decouples the observation path from the behavior path. By actively poisoning the residual signal used for measurement, Bait-and-Recover disrupts the attacker's edit search, while the recovery layer restores clean downstre
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית