כתבה
arXiv cs.CL ·
מיזוג תווים להכרה דיבור מרוב-לשוני: מחקר מערכתי לאורך גודל המודל והתאמה מחדש
Token Merging for Multilingual Speech Recognition: A Systematic Study Across Model Scale and Fine-Tuning
מיזוג תווים משפר את העמידה המחשבתית של הכרה דיבור מרוב-לשוני, עם כמעט אין הפסד בדייקנות. המחקר נעשה על המודל Whisper, וכולל 16 שפות שונות ושלושה גודלי מודל שונים. התוצאות מראות שמיזוג תווים יעיל וכדאי להפחתת עלויות הפעלה של המודל.
תקציר מקורי באנגליתarXiv:2609.13151v1 Announce Type: new Abstract: Leading multilingual speech recognition models like Whisper transcribe diverse, low-resource languages without language-specific training but are computationally expensive to deploy. Token merging mitigates this inefficiency by dynamically combining redundant features, shortening the sequence length during inference without requiring retraining. In this paper, we systematically evaluate token merging on the Whisper model family across sixteen diverse languages and three different model sizes. We also test how token merging interacts with fine-tuning (DoRA) on low-resource languages. Our findings show that merging tokens increases computational efficiency with almost no loss in transcription accuracy across most low-resource languages and mode
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית