כתבה
arXiv cs.LG ·
VoxReason: Listener-Free Evaluation of Source-Grounded Speech Planning Before Synthesis
תקציר מקורי באנגליתarXiv:2609.03203v1 Announce Type: cross Abstract: Expressive speech systems make a decision before any waveform is rendered: how an utterance is delivered. In dialogue agents, narration, and role-conditioned TTS, that hidden planning step sets affect, pitch, energy, rate, pause, emphasis, and stance, yet downstream audio scores rarely reveal whether those choices were licensed by the source record, a source-use failure that occurs before any waveform exists. VoxReason makes that pre-synthesis decision measurable as a listener-free task for source-grounded speech planning. Before synthesis, VoxReason measures whether delivery choices are grounded in cited source records. Systems output a source-cited speaking-plan with evidence citations, and a deterministic verifier checks citation legalit
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית