כתבה
arXiv cs.CL ·
התפשטות-זמן של דובר-יעד ללמידת-מחדש ב-ASR-בעזרת-LLM
Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition
במאמר זה, המחברים מציגים טכניקה חדשה להתפשטות-זמן של דובר-יעד ללמידת-מחדש ב-ASR-בעזרת-LLM. הטכניקה, הקרויה TSU-ASR, מאפשרת למערכת ASR לכתוב את הדיבור של כל הדוברים, למעט אלה שבחרו לא להיות כתובים. המחברים מציגים תוצאות ניסויים על שני סטי נתונים, AMI ו-AliMeeting, ומציעים את הטכניקה כפתרון יעיל למפלטי וידאו-קונפרנסים.
תקציר מקורי באנגליתarXiv:2609.30439v1 Announce Type: new Abstract: We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires an ASR system to transcribe all speakers except the opt-out ones, while still indicating when those speakers are active. As a first step towards tackling this task, we introduce a novel, light-weight Enrollment-Conditioned Gating (ECG) module attachable to a frozen dual-stream speech LLM that enables ASR for new opt-out speakers dynamically during inference, even those who were not seen during initial ECG training phase. Our experiments on both AMI (English) and AliMeeting (Mandarin) data
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית