כתבה
arXiv cs.AI ·
מודלים אקוסטיים רציפים בזמן
Continuous-Time Acoustic Modelling with Neural Controlled Differential Equations
חוקרים הציעו שיטה חדשה למודלים אקוסטיים רציפים בזמן, המשתמשת במשוואות דיפרנציאליות נשלטות על ידי רשתות עצביות. השיטה מאפשרת ייצוג רציף של מידע אקוסטי, ומשפרת את איכות הדיבור הסינתטי.
תקציר מקורי באנגליתarXiv:2609.11725v1 Announce Type: cross Abstract: Text-to-speech (TTS) models commonly address text--speech alignment by expanding phone-level encoder states to frame-level decoder inputs using predicted durations. While this length-regulation step resolves alignment structurally, this use of duration typically changes only where and how often latent states appear, not the values of the states themselves. This paper proposes a continuous-time mechanism for duration-aware acoustic modelling in TTS using neural controlled differential equations (CDEs). We formulate the phone representation as a temporally parameterised control path and use a neural acoustic vector field to produce a continuous-time hidden state whose values evolve with phonetic content and duration-derived timing. The result
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית