כתבה
arXiv cs.LG ·
Finding the Heads and the Neurons Responsible for Network Information Retrieval in Language Models
תקציר מקורי באנגליתarXiv:2610.08200v1 Announce Type: new Abstract: We ask whether specific attention heads, and more finely specific neurons inside those heads, are responsible for recognizing that a language model's context contains network infrastructure information (a hostname paired with its IP address), and whether that responsibility can be validated causally rather than by correlation alone. At the head level the answer is yes, across five models spanning three architecture families: in every model, a small set of heads (1 to 9 out of 128 to 1152 candidates), found by causal ablation screening and tested for selectivity against matched negative and context-free controls, supports a detector with 99.5--100\% held-out accuracy. We then ask whether a head's responsibility concentrates into one neuron or
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית