כתבה
arXiv cs.LG ·
כיצד דגמי שפה שונים בהפצת פעילות ראשי התייחסות תחת דרישה רצפית
How Language Models Differ in Redistributing Attention-Head Activity Under Serial Demand
דגמי שפה שונים בכיצד הם מפצים פעילות ראשי התייחסות תחת דרישה רצפית. חלקם קונצנטרים פעילות בראשי התייחסות החלק האמצעי שלהם, בעוד אחרים פוזרים אותה בראשי התייחסות החלק האחרון שלהם. נמצא כי דגמי Qwen2.5 ו-Llama נבדלים בכיצד הם מפצים פעילות ראשי התייחסות. נמצא גם כי דגמי Qwen2.5-72B ו-Llama-3.1-70B, שיש להם אותו מספר של ראשי תייחסות, נבדלים בכיצד הם מפצים פעילות ראשי התייחסות.
תקציר מקורי באנגליתarXiv:2609.36221v1 Announce Type: new Abstract: The way a model distributes activity over each layer's attention heads offers a coarse view of how it routes information through depth; how this changes with the task is part of what a mechanistic account must explain. Holding prompt length fixed, we vary how many serial steps a task demands and measure, in every layer of 17 open-weight models, whether activity concentrates on a few heads or spreads across many as demand rises. Both occur: in most models, layers just before mid-depth concentrate activity and later layers spread it. Models differ in where and how strongly this happens. The Qwen2.5 base models from 0.5B to 7B, for example, spread less than the average model in every task and concentrate activity in parts of their second half, w
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית