יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

התפשטות-זמן של דובר-יעד ללמידת-מחדש ב-ASR-בעזרת-LLM

Inference-Time Target Speaker Unlearning in LLM-Based Automatic Speech Recognition
במאמר זה, המחברים מציגים טכניקה חדשה להתפשטות-זמן של דובר-יעד ללמידת-מחדש ב-ASR-בעזרת-LLM. הטכניקה, הקרויה TSU-ASR, מאפשרת למערכת ASR לכתוב את הדיבור של כל הדוברים, למעט אלה שבחרו לא להיות כתובים. המחברים מציגים תוצאות ניסויים על שני סטי נתונים, AMI ו-AliMeeting, ומציעים את הטכניקה כפתרון יעיל למפלטי וידאו-קונפרנסים.
תקציר מקורי באנגליתarXiv:2609.30439v1 Announce Type: new Abstract: We introduce target-speaker unlearning ASR (TSU-ASR) task in a fully end-to-end framework for multi-speaker ASR and diarization. Given a multi-speaker utterance and a set of opt-out speakers who do not wish to have their speech transcribed, the task requires an ASR system to transcribe all speakers except the opt-out ones, while still indicating when those speakers are active. As a first step towards tackling this task, we introduce a novel, light-weight Enrollment-Conditioned Gating (ECG) module attachable to a frozen dual-stream speech LLM that enables ASR for new opt-out speakers dynamically during inference, even those who were not seen during initial ECG training phase. Our experiments on both AMI (English) and AliMeeting (Mandarin) data
קרא במקור המקורי