כתבה
arXiv cs.CL ·
Contrastive Projection: Reading Transformer Internals by Differencing Logit Lenses
תקציר מקורי באנגליתarXiv:2609.09902v1 Announce Type: new Abstract: Reading a transformer's internal states in token space is easy to do and hard to trust: a logit lens on a single hidden state is dominated, at intermediate layers, by the generic tokens the model would predict for almost any input. We read the difference instead. Subtracting two closely matched prompts' hidden states and projecting through the unembedding cancels the shared component and surfaces what separates them, an operation equivalent to reading a RepE/ActAdd steering vector through a logit lens. Built into a training-free tracer that reads at every position, sub-layer, and head and averages over designed baselines, it traces a compound- noun MLP->attention chain in Phi-2, confirmed there by activation patching, with the same distinctio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית