כתבה
arXiv cs.LG ·
Towards Isolated Interventions via Almost Orthogonal Features in Language Models
תקציר מקורי באנגליתarXiv:2602.04718v3 Announce Type: replace Abstract: A central premise in mechanistic interpretability is that meaningful concepts in language models are represented by linear features in activation space. For such features to support reliable interventions, manipulating one feature should not substantially alter the effects of others. In practice, however, feature entanglement leads to interference such that localized interventions can have unintended downstream effects. Motivated by the \textit{Independent Causal Mechanisms} principle, we propose to constrain internal features to be almost orthogonal. We argue that this promotes modular representations amenable to causal intervention. We formalize this problem by characterizing the gap between an idealized isolated intervention and its re
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית