כתבה
arXiv cs.AI ·
HoliTok: A Continuous Holistic Tokenization with Robust Dual Capabilities of Speech Generation and Understanding
תקציר מקורי באנגליתarXiv:2605.29948v3 Announce Type: replace-cross Abstract: Unified speech foundation models require a holistic tokenization space that is both learnable by language models and decodable into high-quality waveforms. Existing speech tokenizers, however, often fail to satisfy these requirements simultaneously, leading to increased architectural complexity and more involved training designs. We propose HoliTok, a continuous Holistic speech Tokenization model designed for unified generation-understanding modeling. HoliTok encodes 48~kHz speech into a compact 25~Hz sequence of 128-dimensional latents. It is trained with a progressive strategy that jointly preserves signal-level fidelity, incorporates semantic information, and maintains strong latent learnability. Based on this tokenization, we bu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית