יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

Do LLMs Really Forget? Hidden-State Leakage in Model Unlearning and How to Fix it

מחקר חדש מגלה שמודלי LLMs לא תמיד 'מזניחים' את המידע הסודי, והוא עדיין נשאר במיינורי המודל.
תקציר מקורי באנגליתarXiv:2609.36612v1 Announce Type: new Abstract: Unlearning in large language models (LLMs) is typically evaluated at the output level, where a model appears to suppress sensitive or undesirable content. In this work, we show that such evaluations can create an illusion of forgetting: even when output-level leakage is eliminated, sensitive information can remain encoded in the model's hidden representations. We first provide a theoretical analysis establishing a fundamental separation between output suppression and representational erasure. Specifically, we show that the decoder can be made arbitrarily insensitive to sensitive directions, driving output-level leakage to zero, while the hidden representations retain the underlying information. To empirically validate this phenomenon, we trai
קרא במקור המקורי