כתבה
arXiv cs.LG ·
FASTDIAR: מקדם-ספיקר רציף להפרדת דוברים בזרימה
FASTDIAR: Frame-level speaker encoder for Streaming Diarization
מערכות קונברציונליות באינטרנט דורשות הפרדת דוברים בזמן אמת, שמסוגלת לזרום ולרוץ על מעבד CPU.
תקציר מקורי באנגליתarXiv:2610.02941v1 Announce Type: cross Abstract: Real-time conversational agents require speaker diarization that streams and runs on a CPU. Most systems apply an utterance-level speaker encoder to short, heavily overlapping chunks, which wastes computation and leaves the model optimized for the wrong task. We instead turn a state-of-the-art speaker recognition architecture into a causal frame-level encoder that reads the stream once and emits one embedding every 80~ms from a bounded two-second window of past audio, and pair it with online clustering that gates every update on the self-similarity of the stream. Trained only by distillation from an utterance-level teacher on simulated and out-of-domain mixtures, and evaluated with one fixed set of hyperparameters, the system is the most ac
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית