כתבה
arXiv cs.LG ·
VoxReason: בדיקה-ללא-שומע של תכנון-דיברנסיות-מבוסס-מקור לפני ייצור
VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis
VoxReason מציגה בדיקה-ללא-שומע של תכנון-דיברנסיות-מבוסס-מקור לפני ייצור. המערכת מפיקה תוכנית-דיברנסיות-מבוסס-מקור ומבדק ווריפאי תקין חוקיות התוכנית. המחקר מציג תוצאות טובות של המערכת בבדיקות-ללא-שומע.
תקציר מקורי באנגליתarXiv:2609.03203v2 Announce Type: replace-cross Abstract: Expressive speech systems have to decide how an utterance is delivered before any waveform is rendered. In dialogue agents, narration, and role-conditioned TTS, that planning step sets affect, pitch, energy, rate, pause, emphasis, and stance, yet standard audio metrics rarely show whether those choices were actually licensed by the source record. This leaves a practical evaluation gap: a system may sound plausible while relying on a memorized script instead of the cue that governs delivery. VoxReason casts this pre-synthesis step as a listener-free task for source-grounded speech planning. Systems output a source-cited speaking plan, and a deterministic verifier checks citation legality, slot agreement, unsupported state, schema val
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית