כתבה
arXiv cs.LG ·
מדיניות תשומת הלב-מתגרסת לצורך תרגום דיבור-טקסט בזמן-אמת
Attention-Based Adaptive Policies for Simultaneous Speech-to-Text Translation
במאמר זה, נציגים שתי מדיניות חדשות לתרגום דיבור-טקסט בזמן-אמת, המשתמשות במנגנון תשומת הלב-מקביל. המדיניות החדשות, RFAP ו-DCAP, מאפשרות למודלי תרגום דיבור-טקסט להיות משולבים בסקנרי זרימה בלי צורך באימון מחדש.
תקציר מקורי באנגליתarXiv:2609.30839v1 Announce Type: new Abstract: Simultaneous speech-to-text translation (Simul-S2TT) consists of generating partial translations while the incoming audio frames are processed by the system. However, the streaming nature of this setup creates the challenge of deciding the best moment to perform an accurate translation while minimizing the delay. To address this challenge, we utilize the cross-attention mechanism of the encoder-decoder architecture to find the right alignment between the input speech frames and the target text tokens. In this paper, we propose the Recent Frame Attention Policy (RFAP) and the Dual-Condition Attention Policy (DCAP) that allow offline trained speech-to-text translation models to be used in streaming scenarios without requiring additional trainin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית