יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

DuoTok: תקן-מקורי דו-עקבה לטוקניזציה של מוזיקה לשירה-ליווי

DuoTok: Source-Aware Dual-Track Music Tokenization for Vocal-Accompaniment Generation
DuoTok הוא תקן-מקורי דו-עקבה לטוקניזציה של מוזיקה לשירה-ליווי, המשמר את האיכות האקוסטית ומקסים את העבירות-ספקטרליות. הוא משתמש באלגוריתם של דיס-אנטנגלמנט, ומקסים את העבירות-ספקטרליות ואת האיכות האקוסטית. DuoTok הוא תקן-מקורי דו-עקבה לטוקניזציה של מוזיקה לשירה-ליווי, המשמר את האיכות האקוסטית ומקסים את העבירות-ספקטרליות.
תקציר מקורי באנגליתarXiv:2511.20224v3 Announce Type: replace-cross Abstract: Multi-track music generation requires tokens that preserve acoustic fidelity, support sequence modeling, and maintain cross-track structure. Reconstruction-oriented codecs retain acoustic detail but are difficult to model, while semantic tokenizers may sacrifice fidelity or cross-track alignment. We present DuoTok, a source-aware dual-track music tokenizer for vocal-accompaniment generation based on staged disentanglement. DuoTok first learns a semantic audio representation through self-supervised pretraining, then shapes source-aware structure using feature replacement noise and multi-task supervision: spectral reconstruction, music source separation regularization, and an ASR head for lyric alignment. It freezes the encoder and le
קרא במקור המקורי