כתבה
arXiv cs.LG ·
חשיפת פרומפטים נשכחים ממודלים
Extracting Forgotten Prompts from Targeted Unlearned Models
חוקרים גילו פרצת אבטחה חדשה המאפשרת לחשוף פרומפטים נשכחים ממודלים שעברו תהליך unlearning. התקפה זו, Targeted Active Search (TAS), מסוגלת לשחזר עד 95% מהפרומפטים הנשכחים.
תקציר מקורי באנגליתarXiv:2609.03662v1 Announce Type: new Abstract: Recent unlearning methods (e.g. NPO, DPO, LUNAR) make use of refusal alignment to suppress forgotten data. However, it has been shown that refusal responses might leave traces of unlearning, and recent attacks have been able to successfully recover some of the unlearned knowledge. In this paper, we uncover a new vulnerability. Existing attacks typically assume that the forgotten prompts are already known to the adversary and focus on recovering their answers. However, we show that the forgotten prompts themselves can be extracted by using the retained data and black-box access to the model. Our attack, Targeted Active Search (TAS), first identifies the forgotten entities by constructing canonical templates and entity pool, and selectively que
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית