יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

גאומטריה משותפת כאבן רוזטה: התאמה מודלית-מודלית ללא נתוני זוגות

Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data
במאמר זה, נראה שאפשר לבצע התאמה מודלית-מודלית ללא נתוני זוגות. התוצאות יכולות לאפשר יצירת תמונות מטקסט ללא נתוני זוגות.
תקציר מקורי באנגליתarXiv:2610.09411v2 Announce Type: replace Abstract: Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation geometry. But then, do we even need paired examples for cross-modal alignment? Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment. Our simple Wasserstein Procrustes method with a coarse geometric initialization aligns two disjoint embedding sets by estimating a single orthogonal map without seeing any pairs. Across datasets, modalities, and unimodal models, we show that we can consistentl
קרא במקור המקורי