כתבה
arXiv cs.CL ·
VOSSA: אופטימיזציה של תבנית קול לאופטימיזציה של תבנית קול לארכיטקטורות זרימה
VOSSA: Voiceprint Optimization for Streaming Speech Architectures
VOSSA (Voiceprint Optimization for Streaming Speech Architectures) היא תצורה של נציגות דיבור שמאפשרת אופטימיזציה של תבנית קול לארכיטקטורות זרימה. VOSSA משתמשת באגרגטציה של statistic pooling ומאפשרת תרגום דיבור בזמן אמת.
תקציר מקורי באנגליתarXiv:2609.38887v1 Announce Type: cross Abstract: Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effective for speaker discrimination, these embeddings are trained to remain stable across phonetic and prosodic variations within-speaker, which may conflict with frame-level acoustic generation in streaming constraints. To address this issue, we propose VOSSA (Voiceprint Optimization for Streaming Speech Architectures), a speaker representation framework that extracts speaker information from intermediate content encoder layers and aggregates using attentive statistics pooling. The embedding is trained jointly with VC objectives, removing the need for a separate speaker encoder. Across six datasets
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית