כתבה
arXiv cs.LG ·
Repeated-Token Counting Reveals a Dissociation Between Representations and Outputs
תקציר מקורי באנגליתarXiv:2605.09239v2 Announce Type: replace-cross Abstract: Large language models fail at counting how many times a word repeats in a list, even though they perform well on far harder reasoning tasks. These failures are commonly attributed to limitations in internal count tracking. We show this attribution is wrong. Linear probes on the residual stream decode the correct count with near-perfect accuracy at every post-embedding layer and they do so even at the exact layers where the wrong answer crystallizes in the output. Attention patterns show no evidence of collapse over repeated tokens and tokenization artifacts account for none of the failure. Instead, a multi-layer perceptron (MLP) block at roughly 85--93\% network depth overwrites the correctly-encoded count with a fixed wrong answer.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית