כתבה
arXiv cs.CL ·
הנחיית מותאמת לדובר ב-TS-ASR
Asymmetric Classifier-Free Guidance for Target-Speaker ASR
אנו מציגים הנחיית מותאמת לדובר להכרה שפתית יעד (TS-ASR) שמשפרת את תיקון השגיאות ב-21.8%.
תקציר מקורי באנגליתarXiv:2609.30476v1 Announce Type: cross Abstract: Target-speaker automatic speech recognition (TS-ASR) must identify and transcribe a desired speaker under varying overlap and noise conditions. These changes alter the acoustic evidence for the target speaker in the speech mixture, motivating inference-time calibration of speaker conditioning. We introduce asymmetric classifier-free guidance (CFG) for TS-ASR using Whisper: the speaker-conditioned branch predicts the target transcript, while the speaker-unconditioned branch predicts serialized multi-speaker transcripts. CFG adjusts the contribution of speaker conditioning during decoding through a single guidance scale. We select a global guidance scale on target-domain development data and train a lightweight encoder-based predictor to adju
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית