כתבה
arXiv cs.CL ·
עברת ראשים: דגלי שפה עוקבים אחר דוברים עם תשומת לבם של הגוף הכתוב שלהם, ודירוג תשומת לב-מסה מחזיר ראש אחר
Inherited Heads: Audio language models track speakers with their text backbone's attention, and an attention-mass ranking retrieves a different set
דגלי שפה עוקבים אחר דוברים עם תשומת לבם של הגוף הכתוב שלהם, ודירוג תשומת לב-מסה מחזיר ראש אחר. נמצא כי ראשי הגוף הכתוב של הדגל עוזרים לדגל לעקוב אחר הדובר. נמצא גם כי דירוג ראשי הגוף הכתוב על פי תשומת לבם לפסקא השאלה נותן תוצאות טובות יותר מדירוג ראשי הגוף הכתוב על פי תשומת לבם המשונה.
תקציר מקורי באנגליתarXiv:2609.14174v1 Announce Type: cross Abstract: Asked to describe what one of six speakers in a recording talks about, audio language models describe the right one on 6 to 16% of trials, below the 16.7% a guess would give. Adding a fixed bias to the attention logits of a hundred heads, under a tenth of the model's and with no training, redirects the description to whichever speaker we choose, on 90.7% to 99.0% of trials. Those heads are largely not specific to audio. Rank the text-only language model an audio model was built from, or a released model of the same family, on a written version of the task, take its top hundred heads, and carry them over unchanged: they redirect the audio model on 80.8% to 95.0% of trials, with nothing about audio entering the selection. The audio and text h
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית