כתבה
arXiv cs.LG ·
גאומטריה משותפת כאבן רוזטה
Shared Geometry As A Rosetta Stone: Cross-Modal Alignment Without Paired Data
חוקרים מצאו דרך ליישר כוחות בין מודלים רב-מודאליים ללא צורך בנתונים מזוגים. השיטה החדשה משתמשת בגאומטריה משותפת כדי ליצור התאמה בין ייצוגים שונים.
תקציר מקורי באנגליתarXiv:2610.09411v1 Announce Type: new Abstract: Multimodal representations enable zero-shot classification and retrieval, but aligning independently trained models usually requires large amounts of paired data. Yet, the Platonic Representation Hypothesis suggests that models trained on different modalities may converge spontaneously toward a shared representation geometry. But then, do we even need paired examples for cross-modal alignment? Remarkably, we show that paired examples are unnecessary for coarse cross-modal alignment. Our simple Wasserstein Procrustes method with a coarse geometric initialization aligns two disjoint embedding sets by estimating a single orthogonal map without seeing any pairs. Across datasets, modalities, and unimodal models, we show that we can consistently al
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית