כתבה
arXiv cs.AI ·
CoJEPA: Combining Contrastive Learning and JEPA for Global-Local Music Representations
תקציר מקורי באנגליתarXiv:2608.30974v2 Announce Type: replace-cross Abstract: Joint-Embedding Predictive Architecture (JEPA) has shown strong performance in learning rich representations through self-supervised prediction in latent space. However, it typically relies on teacher--student architecture with an EMA to stabilise training, and can tend to yield uninformative representations. Contrastive learning is stable to train and produces strong global representations, but remains limited on local tasks by the global nature of its objective. In this work, we combine both into CoJEPA: a single shared backbone jointly trained with a JEPA objective on masked sequence tokens and a contrastive objective on the class token. The contrastive gradient provides stability, removing the need for an EMA teacher entirely, w
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית