כתבה
arXiv cs.LG ·
Certified Mechanistic Edits: Behavioral Guarantees for Skill Removal and Preservation
תקציר מקורי באנגליתarXiv:2610.03502v1 Announce Type: new Abstract: Mechanistic edits (ablations, weight edits, activation steering) are the standard tools for unlearning a harmful capability from a neural network while preserving useful ones. Current approaches validate their effects only by testing, which can never cover an entire continuous region of inputs. Prior work at the interpretability-verification boundary certifies descriptions of a model: what a circuit computes, or whether it faithfully explains the whole. We instead certify the behavioral effect of an edit: that disabling a circuit removes one skill and provably preserves another, for every input in a region; a feature non-interference guarantee in the information-flow-security sense. We demonstrate such certified edits from toy ReLU networks u
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית