כתבה
arXiv cs.LG ·
Storage Is Not Strategy: State-Conditioned Support Control for LLM Unlearning
תקציר מקורי באנגליתarXiv:2609.37858v2 Announce Type: replace Abstract: Many localized large language model (LLM) unlearning methods select a small parameter subset from a localization signal and keep it fixed during optimization. The parameters most associated with a target, however, need not be the best ones to update, and candidate interventions can change value as optimization proceeds. In a controlled experiment, a storage-localization score reaches an area under the receiver operating characteristic curve (AUROC) of 0.981, yet storage identity agrees with the better intervention on only 17/36 targets, while low-rank adaptation (LoRA) wins 35/36. We introduce Intervention Score, which ranks editable groups by the predicted effect of the actual unlearning update while accounting for collateral damage, and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית