כתבה
arXiv cs.LG ·
אימות אי-למידה של מושגים במודלים דיפוזיים
Certifying Concept Unlearning in Text-to-Image Diffusion Models
חוקרים פיתחו כלי לאימות אי-למידה של מושגים במודלים דיפוזיים ליצירת תמונות מטקסט. הכלי מספק הבטחות בטוחות עם שגיאה מוגבלת על הדליפה של מושגים. המחקר בדק שלוש קטגוריות עיקריות: תוכן NSFW, סגנונות אמנותיים וזהויות ידוענים.
תקציר מקורי באנגליתarXiv:2609.12163v1 Announce Type: new Abstract: Existing evaluations of concept unlearning in text-to-image (T2I) diffusion models primarily rely on attack success rates obtained through automated adversarial prompt search. However, these metrics provide only empirical evidence over a finite set of queries and leave residual leakage over the broader prompt space largely unquantified. This limitation can lead to overestimating unlearning effectiveness and underestimating safety risks. To address this gap, we introduce a novel certification framework for T2I concept unlearning that provides high-confidence guarantees with bounded error on residual concept leakage. Our approach combines statistical certification with worst-case analysis along concept-relevant embedding directions to derive ex
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית