כתבה
arXiv cs.LG ·
TontaubeV1: Streaming Text-to-Speech with Hierarchical Codec Modeling and Bounded Context
תקציר מקורי באנגליתarXiv:2609.08703v1 Announce Type: cross Abstract: Text-to-speech systems often face a trade-off between natural prosody and efficient inference: higher perceptual quality typically comes at increased computational cost and latency. We present TontaubeV1, a model that preserves natural prosody while enabling streaming from a single consumer GPU. Speech is encoded by the hierarchical DualCodec representation at 12.5 Hz, which separates a semantic stream from successive acoustic refinements. Our design assumes that prosodic structure is largely established when the semantic stream is generated, and allocates capacity accordingly: a Qwen3-1.7B-derived transformer predicts that stream and thereby the utterance duration, while three progressively smaller Qwen3-0.6B-derived transformers each add
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית