יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

זיהוי דיבור אודיו-ויזואלי בתנאים של מחסור במשאבים

Bootstrapping Audiovisual Speech Recognition in Zero-AV-Resource Scenarios
חוקרים מצליחים לשפר זיהוי דיבור אודיו-ויזואלי בשפות עם משאבים מוגבלים. הם משתמשים בנתונים ויזואליים סינתטיים כדי לאמ� את המודל. התוצאות מראות שיפור משמעותי ביכולת הזיהוי.
תקציר מקורי באנגליתarXiv:2603.08249v2 Announce Type: replace-cross Abstract: Audiovisual speech recognition (AVSR) combines acoustic and visual cues to improve transcription robustness under challenging conditions but remains out of reach for most under-resourced languages due to the lack of labeled video corpora for training. Synthetic visual data have been shown to be an effective augmentation strategy for addressing AV data scarcity. However, a more challenging scenario arises for languages such as Catalan, where no real audiovisual data are available for training. In this study, we investigate whether AVSR can be bootstrapped in such a zero-AV-resource setting, using synthetic visual data as the sole source of visual supervision. We synthesize over 700 hours of talking-head video and fine-tune a pre-trai
קרא במקור המקורי