כתבה
arXiv cs.AI ·
משולש הקשב במודלים של אודיו-וידאו
The Attention Triangle in Audio-Video Models
חוקרים את מנגנון הקשב המשולש במודלים של אודיו-וידאו, המאפשר תקשורת בין טקסט, קול ותמונה. הניתוח מראה כי הקשב הדו-כיווני בין האודיו לווידאו יכול לגרום לדליפת מידע סמנטי. המחקר מציע כלים לאבחון ותיקון התופעה.
תקציר מקורי באנגליתarXiv:2609.03586v1 Announce Type: new Abstract: Audio-video diffusion models rely on cross-modal attention to coordinate text, sound, and visual content, yet this same mechanism can introduce subtle and systematic semantic leakage. We study these models by probing and analyzing the ``attention triangle,'' comprising the three cross-attention edges connecting the text, audio, and video streams, and examine how semantic information is routed across modalities during generation. Our analysis reveals that routing along the audio-video edge is bidirectional: audio can influence video generation, while video can influence audio generation. This edge is shaped by biases encoded in the model's parameters and emerges as a major contributor to leakage: when prompts are in tension with learned priors
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית