כתבה
arXiv cs.LG ·
WhisTLE: תאדפטציה עמוקה, רק טקסט, למודלי הבינה המלאכותית להכרה בתמונות
WhisTLE: Deeply Supervised, Text-Only Domain Adaptation for Small Pretrained Speech Recognition Transformers
WhisTLE היא שיטת תאדפטציה עמוקה, רק טקסט, למודלי הבינה המלאכותית להכרה בתמונות. היא משתמשת ב-VAE כדי ללמוד את הקודקס של המודל ולאדפט על-ידי טקסט. WhisTLE עוזרת לשפר את ה-WER ב-88% מהמקרים.
תקציר מקורי באנגליתarXiv:2509.10452v3 Announce Type: replace-cross Abstract: Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance. In many real-world settings, collecting speech data is impractical, necessitating text-only adaptation. We propose WhisTLE, a deeply supervised, text-only adaptation method for pretrained encoder-decoder ASR models. WhisTLE trains a variational autoencoder (VAE) to model encoder outputs from text and fine-tunes the decoder using the learned text-to-latent encoder, optionally combined with text-to-speech (TTS) adaptation. At inference, the original encoder is restored, incurring no extra runtime cost. Across four datasets and four ASR models, WhisTLE alone helps in 28 of 32 (88%) cases, with an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית