כתבה
arXiv cs.LG ·
Decoder-Side Semantic Conditioning for Low-Bitrate Neural Speech Compression
תקציר מקורי באנגליתarXiv:2512.21653v2 Announce Type: replace-cross Abstract: Speech codecs are usually optimized for waveform fidelity, allocating bits to acoustic detail that can be inferred from linguistic structure. This leads to inefficient compression and degraded recognition performance. We propose SemDAC, a semantic-aware neural speech codec that adds hierarchical semantic conditioning to residual vector quantization (RVQ). The first RVQ quantizer is distilled from HuBERT features to produce semantic tokens capturing phonetic content, while later quantizers encode residual acoustics. The decoder is conditioned on semantic tokens via feature-wise linear modulation (FiLM), steering reconstruction toward information not explained by semantic abstraction. At 0.95 kbps, SemDAC matches or surpasses a 2.5 kb
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית