כתבה
arXiv cs.LG ·
Nullify: שליטה באקטיבציה במרחב האפס להשמדת-ללא-אימון של LLM
Nullify: Null-Space Activation Steering for Training-Free LLM Unlearning
אנו מציגים Nullify, שיטה ללא-אימון להשמדת-בוחרת של נתונים פרטיים ב-LLMs. Nullify משתמשת בווקטורי שליטה במהלך הערכה כדי לפנות את האקטיבציות הקשורות לפרטיות לעבר התשובות המזוהרות, בעוד שהאקטיבציות השאריות נשארות ללא שינוי. Nullify נבחנה על TOFU ו-MUSE והוכיחה כי היא עוקבת או עולה על הבסיסים המוכרים באיכות השמדה, בעוד שהיא שומרת על תכונות המודל בעיקר.
תקציר מקורי באנגליתarXiv:2610.10655v1 Announce Type: new Abstract: Large Language Models (LLMs) inevitably internalize substantial amounts of sensitive or private information during pre-training, while LLM unlearning aims to selectively erase specific knowledge to prevent privacy leakage with minimal loss of model utility. However, existing methods struggle to balance forget quality with utility, and typically incur substantial computational costs due to parameter fine-tuning. To address this, we propose Nullify, a training-free, non-destructive activation steering method for LLM unlearning. Nullify employs steering vectors during inference to redirect privacy-related activations away from their memorized answers, while satisfying a null-space constraint that leaves retained-query activations essentially una
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית