כתבה
arXiv cs.AI ·
Characterize Then Distill: Mechanistic Reasoning in Large Output Spaces
תקציר מקורי באנגליתarXiv:2606.06840v2 Announce Type: replace-cross Abstract: Reasoning-trained language models can perform, zero-shot, multi-label tasks that require selecting a small set of relevant labels from a universe of thousands to hundreds of thousands of candidates. We ask how they do it mechanistically, and whether the mechanism can be distilled. We make the question measurable by treating each decision as a token-level event scored by the model's own decision margin: the token that picks a coarse region of the label space, the tokens that pick a label within it, and the token where the output departs from a close alternative (a near-miss) named earlier in the reasoning. Attribution, exact mean-ablation, knock-in into another example's context, and a null calibration that discounts generic heads th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית