כתבה
arXiv cs.LG ·
Contrastive Representation Shaping for LLM Unlearning
תקציר מקורי באנגליתarXiv:2601.22028v2 Announce Type: replace Abstract: Most LLM unlearning methods aim to approximate retrain-from-scratch behaviors with minimal distribution shift, often via alignment-style objectives defined in the prediction space. While effective at reducing forgotten content generation, such approaches may act as suppression: forgotten concepts can persist in representations and remain entangled with retained knowledge. We introduce CLReg, a contrastive representation regularizer that identifies forget features while pushing them away from retain features, reducing forget--retain interference while empirically preserving the scale and shape of retain features. As light motivation for the mechanism, we provide a one-step analysis showing that CLReg decreases a simple entanglement proxy i
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית