כתבה
arXiv cs.CL ·
HPRO: אופטימיזציה היררכית ודינאמית להשבחת פרסות על-פי קביעת טעמים לטקסט-לדיבור
HPRO: Hierarchical Progressive Reward Optimization via Preference Extraction for Emotional Text-to-Speech
HPRO היא פרקטיקה של אופטימיזציה היררכית ודינאמית שמטרתה לשפר את הביטוי הרגשי של טקסט-לדיבור. הפרקטיקה נועדה לפצל את האופטימיזציה לפי טעמים שונים, כדי לשפר את הביטוי הרגשי. הפרקטיקה נבחנה באמצעות ניסויים שהראו תוצאות טובות. הקוד והדוגמאות הקוליות של HPRO זמינים באתר https://xxh333.github.io/hpro-demo/
תקציר מקורי באנגליתarXiv:2606.28249v2 Announce Type: replace-cross Abstract: Recently, Large Language Model (LLM)-based Text-to-Speech (TTS) models have achieved remarkable naturalness. However, the standard Supervised Fine-Tuning paradigm often converges to statistically averaged prosody, limiting emotional expressiveness. While preference-driven optimization offers a promising alternative, existing approaches suffer from two structural mismatches: information conflict, where content and emotion in a shared latent space produce conflicting gradients, leading to reward hacking and semantic degradation; and scale gap, where sparse sentence-level rewards struggle to guide dense frame-level generation. To overcome these challenges, we propose HPRO, a hierarchical progressive reward optimization framework. Withi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית