כתבה
arXiv cs.CL ·
האם ייצוג מאולף לפרוזודיה עוזר מעבר לשילוב מאולף?
Does a prosody-trained representation help beyond trainable fusion? A parameter-matched study with frozen HuBERT
חוקרים בדקו האם ייצוג מאולף לפרוזודיה יכול לשפר את רמת הדיוק של הכרת דיבור. הם השוו בין מודלים שונים ומצאו שהייצוג המאולף לא הביא לשיפור משמעותי.
תקציר מקורי באנגליתarXiv:2609.36754v1 Announce Type: cross Abstract: Explicit prosodic cues may help automatic speech recognition (ASR) of spontaneous speech, but auxiliary representations typically require additional trainable components, making it unclear whether gains come from the auxiliary information or the fusion mechanism. We address this using a frozen HuBERT backbone and a 64-dimensional representation trained to predict log F0, voicing, Delta log F0, log energy, and spectral tilt. We compare a frozen-backbone recognizer (Baseline), trainable fusion with zero auxiliary input (Null), and the same fusion supplied with the learned representation (Learned). Across Buckeye, Switchboard, and AMI IHM, Null reduces WER by 0.71-1.45 points over Baseline, whereas Learned differs from Null by +0.07, -0.09, an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית