יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מחקר מקיף על ייצוגי תוכן לסינתזה של דיבור

A Comprehensive Study of Content Representations for Speech Synthesis
חוקרים בדקו ייצוגי תוכן שונים לסינתזה של דיבור, ומצאו שייצוגים מסוימים מצליחים להפריד בין זהות הדובר לתוכן הדיבור. התוצאות מראות שההפרדה תלויה לא רק בהדרכה, אלא גם ביכולת הייצוג.
תקציר מקורי באנגליתarXiv:2609.30975v1 Announce Type: cross Abstract: Speech content representations are central to voice conversion, speech-to-speech translation, and multimodal language models, yet they are rarely compared under a common generative framework that directly measures what each representation contains. We address this by training a generative model conditioned solely on each representation and evaluating the generated audio along the content, speaker identity, and prosody axes. Across SSL features, supervised tokens, posteriorgrams, and neural audio codecs, we find two distinct regimes: representations that nearly reconstruct the original audio, and representations that effectively disentangle speaker identity. These results show that disentanglement depends not on supervision alone, but on the
קרא במקור המקורי