יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

עברת ראש: דגלי שפה קוליים עוקבים אחר דוברים עם תשומת לבם של הראשיות של הגוף הכתוב

Inherited Heads: Audio language models track speakers with their text backbone's attention, and an attention-mass ranking retrieves a different set
דגלי שפה קוליים עוקבים אחר דוברים עם תשומת לבם של הראשיות של הגוף הכתוב. המחקר מציע דרך לשפר את היכולת של הדגלים לזהות את הדובר הנכון.
תקציר מקורי באנגליתarXiv:2609.14174v1 Announce Type: new Abstract: Asked to describe what one of six speakers in a recording talks about, audio language models describe the right one on 6 to 16% of trials, below the 16.7% a guess would give. Adding a fixed bias to the attention logits of a hundred heads, under a tenth of the model's and with no training, redirects the description to whichever speaker we choose, on 90.7% to 99.0% of trials. Those heads are largely not specific to audio. Rank the text-only language model an audio model was built from, or a released model of the same family, on a written version of the task, take its top hundred heads, and carry them over unchanged: they redirect the audio model on 80.8% to 95.0% of trials, with nothing about audio entering the selection. The audio and text hea
קרא במקור המקורי