כתבה
arXiv cs.LG ·
MUtE: A Dual Framework for Concept Erasure and Counterfactual Interventions
תקציר מקורי באנגליתarXiv:2609.11253v1 Announce Type: new Abstract: Erasing concept-specific information from representations has been proven useful for mitigating bias or interpreting model decisions. The joint objective is to transform the original representations such that the target concept becomes unpredictable, while maximally preserving concept-unrelated information. In this work, we revisit the optimal bounds of concept erasure to derive a novel class of erasure functions that naturally induce a deterministic, dual counterfactual mapping. Bridging the gap between theoretical optimality and practical representation learning, we design an implementation that imposes a translational bias on counterfactual trajectories - a constraint that aligns with how many concepts geometrically manifest in modern lang
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית