כתבה
arXiv cs.LG ·
Weights Read and Write Features: Scalable Parameter Decomposition Grounded in Activation Space
תקציר מקורי באנגליתarXiv:2609.37731v1 Announce Type: new Abstract: Activation space and parameter space provide complementary views of model computation. Activations represent information, while weights read, transform, and write that information. Yet existing interpretability methods largely study the two spaces separately, leaving the connection between represented information and parameter-level computation underexplored. We introduce Activation-Supported Parameter Decomposition (ASPD), which jointly decomposes activation and parameter spaces and grounds each learned weight component in the activation features it reads or writes. This grounding constrains otherwise non-unique parameter decompositions using the model's internal activations, while an internal reconstruction objective provides a local learni
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית