כתבה
arXiv cs.CL ·
Content is What Remains: Invariant Speech Tokenization from Parallel Utterances
תקציר מקורי באנגליתarXiv:2607.19033v1 Announce Type: new Abstract: Discrete speech tokenizers aim to disentangle semantic from acoustic information, yet targets from self-supervised learning (SSL) models like HuBERT retain non-linguistic variation: speaker identity, prosody, and channel conditions leak into the tokens, inflating entropy. Our key insight is that when enough speakers utter the same words under varying conditions, linguistic content is the only shared factor. We propose PINT (Parallel INvariant Tokenization), which fine-tunes an SSL encoder with alignment losses across parallel utterances and augmentations to distill this shared residual. PINT collapses identical words onto consistent token sequences, drastically reducing conditional entropy. Unlike ASR text, PINT tokens preserve frame-level te
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית