כתבה
arXiv cs.LG ·
Signatures of Steerability in Activation Space of Language Models
תקציר מקורי באנגליתarXiv:2609.14151v1 Announce Type: new Abstract: Steering language models using a set of contrastive representations has been a canonical and computationally efficient method for controlling model behavior. Despite this success in controlling certain model behaviors, the effectiveness of activation steering varies markedly across concepts; the generalization properties of steering vectors are often considered a function of the dataset used to construct them. We make this dataset-dependence claim more rigorous and show that simple separation metrics strongly correlate with the downstream steerability of language models across diverse settings, even after controlling for layers and dataset effects. Beyond prediction, we provide evidence from a synthetic superposition experiment that separatio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית