כתבה
arXiv cs.AI ·
LoSATok: מטוקניזציה סמנטית-אקוסטית בממדים נמוכים
LoSATok: Low-Dimensional Semantic-Acoustic Tokenizer for Cross-Domain Audio Understanding and Generation
LoSATok הוא מטוקניזציה סמנטית-אקוסטית בממדים נמוכים להבנה ויצירה של אודיו. הוא משתמש בבקבוק סמנטי כדי לדחוס מאפיינים סמנטיים ל-128 ממדים, ומאפשר הבנה ויצירה משופרות של אודיו.
תקציר מקורי באנגליתarXiv:2605.27840v2 Announce Type: replace-cross Abstract: Audio tokenizers are fundamental to unifying audio understanding and generation. Understanding requires high-level semantics, while generation demands semantic and acoustic details. Existing unified tokenizers jointly encode both in high-dimensional continuous latents, which increases the modeling burden of Diffusion Transformers (DiTs) for generation. We propose LoSATok, a low-dimensional audio tokenizer for cross-domain audio understanding and generation. Motivated by the observation that 1280-dimensional semantic encoder features are compressible, we introduce a Semantic Bottleneck that compresses them into 128 dimensions, regularized by the proposed time-relation loss for temporal feature consistency. We further design a dual-le
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית