כתבה
arXiv cs.CL ·
צורה נבדלת במרחב התנועה הפריפריאלי של דגמי שפה
Emergent Structure in the Marginal Attention Space of Language Models
נמצאה צורה נבדלת במרחב התנועה הפריפריאלי של דגמי שפה. נמצא קשר בין התנהגות המודלים לג'קוביאן של הרשת. קוד זמין ב-GitHub.
תקציר מקורי באנגליתarXiv:2610.03109v1 Announce Type: new Abstract: While representation similarity across independently trained language models is well-documented, how internal mechanics such as attention behave across models remains far less characterized. Inspired by this gap, we examine the structure of post-softmax attention weights by marginalizing over query positions, mapping them into a joint token-head "marginal attention space". Evaluating across 60+ diverse LLMs, we find that different properties emerge when reducing this space along its token and head axes. When reduced token-wise, marginal attention yields a text-intrinsic signal robustly conserved across models. To explain this property, we empirically connect marginal attention to the input-output Jacobian of the network, and prove theoretical
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית