יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

DiTAR+: חקירה כפולה לאופטימיזציה רב-מודלית לשיפור יצירת דיבור

DiTAR+: Dual Optimization for Robust Autoregressive Diffusion Speech Synthesis
DiTAR+ היא פרקטיקה חדשה לשיפור יצירת דיבור, המשלבת שני רעיונות: סימולציה של רצף זמן מורחב והסתרת פרטים אקוסטיים. זה יוצר שיפורים משמעותיים באיכות הדיבור, כולל קיטוע טעויות פרונונציה והיפוכות של תוכן.
תקציר מקורי באנגליתarXiv:2609.13909v1 Announce Type: cross Abstract: Continuous-latent Autoregressive Diffusion Transformer (AR-DiT) models have demonstrated immense potential in zero-shot speech generation. However, they still suffer from limited decoding stability when synthesizing long utterances or complex linguistic structures. This instability primarily stems from a restricted historical receptive field and an acoustic inertia dependency within the diffusion decoder, which causes the model to ignore semantic conditions. To address these challenges, we propose DiTAR+, a dual-optimization framework. First, we introduce Dilated Context Sampling to expand the macro-level historical receptive field without violating physical temporal continuity, thereby preventing cumulative error propagation. Second, we pr
קרא במקור המקורי