כתבה
arXiv cs.LG ·
Local Multimodal Music Alignment from Global Supervision
תקציר מקורי באנגליתarXiv:2607.10023v2 Announce Type: replace-cross Abstract: Understanding music requires understanding localized relationships across data modalities, e.g., how time in performance audio maps onto position in a score image. Yet supervision for such local correspondences is difficult to obtain-in practice, we often only have access to coarser global supervision like paired segments of audio and images. To address this gap, we propose FuSiLi (Fused Sinkhorn-Localized Similarity), a similarity score for multimodal contrastive learning operating directly on local image patch and audio frame features via Sinkhorn-based soft alignment. We show that FuSiLi (i) effectively learns local relationships, (ii) requires only global supervision, and (iii) retains the global alignment capabilities of conven
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית