כתבה
arXiv cs.CL ·
Train the Model, Not the Reader: Decodability Supervision for Verifiable Activation Explanations
תקציר מקורי באנגליתarXiv:2607.20379v1 Announce Type: cross Abstract: Natural-language autoencoders score explanations of hidden activations by reconstruction: an explanation is deemed faithful if the activation can be regenerated from it. The test is structurally insensitive to individual false claims: if flipping a claim does not change the reconstruction, the claim is never penalized. We show the test is passed in two ways, neither faithful. On a released Qwen-2.5-7B verbalizer, explanations reconstruct well above chance while ~2% of specific claims are reconstruction-dependent, so the score tracks gist, not specific facts. Under exact synthetic ground truth, the standard recipe develops co-adapted private codes (false wording the reconstruction depends on) in 5/5 runs, and fixes that leave the target mode
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית