כתבה
arXiv cs.AI ·
עבירה על העקבה: ייצוג-רמת-מידע-בוחר-להשתק
Neuralyzing the Trace: Selective Representation-Level Unlearning with Contrastive Sparse Autoencoders
מאמר זה עוסק בהשתקת מודלים מולטי-לשוניים על ידי ייצוג-רמת-מידע-בוחר-להשתק. המחברים הציגו כלי חדש, SCALPEL, שמסוגל ללמוד ייצוגים יותר נבחרים. הם הציגו תאורטיקל ואתגרים על ידי ניסויים על מודלי Qwen, Llama ו-Gemma.
תקציר מקורי באנגליתarXiv:2609.31056v1 Announce Type: new Abstract: Machine unlearning aims to remove targeted information while preserving a model's other abilities. In realistic settings, such as privacy requests under the EU GDPR, the target may be narrow, for example information associated with a single person. Behavioral forgetting alone may be insufficient, motivating interventions directly on internal representations. However, standard mechanistic-interpretability extractors are poorly selective for such targets. We identify an energy bias in reconstruction-based extraction, which favors dominant background structure over low-energy target-specific components. We introduce SCALPEL, a contrastive sparse autoencoder designed to learn more selective forget features. We show theoretically that contrastive
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית