כתבה
arXiv cs.AI ·
Learning What to Forget: Distributional Unlearning for LLM Representation Spaces
תקציר מקורי באנגליתarXiv:2609.38929v1 Announce Type: new Abstract: Machine learning systems increasingly face the need to remove the influence of entire data domains, such as toxic language, harmful behavior, or topical content, rather than isolated records. Recent work formalizes this problem as \emph{distributional unlearning}: selecting a subset of a forget domain whose removal moves the training distribution away from an unwanted population while preserving proximity to the desired one. However, existing analyses often impose parametric assumptions to obtain tractable selection rules. These assumptions may be poorly suited to high-dimensional language-model representations. We introduce \textsc{Mamushi}, a framework for non-parametric distributional unlearning that ranks forget examples using a probabili
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית