כתבה
arXiv cs.AI ·
גיאומטריה אחת, תוצאות שונות
One Geometry, Different Outcomes: Readout-Dependent Effects of the Modality Gap in Vision-Language Models
מודלים ויז'ואליים-לשוניים לומדים מרחבי הטבעות משותפים על ידי הצמדת זוגות תמונה-טקסט, אך הייצוגים שלהם נשארים מופרדים על ידי פער מודאלי. המחקר מציע הסבר גיאומטרי אחיד להשפעות המשתנות של הפער הזה.
תקציר מקורי באנגליתarXiv:2609.36101v1 Announce Type: cross Abstract: Contrastive vision-language models learn shared embedding spaces by aligning matched image-text pairs, yet their representations remain separated by a modality gap. Prior work reports divergent effects of modifying this gap: reducing it can improve zero-shot classification and cross-modal alignment, whereas removing gap-related structure can degrade image-text retrieval. In this paper, we provide a unified geometric explanation for these task-dependent effects. Across CLIP and SigLIP encoders, we find that a single dominant direction captures 94.4-99.9% of the squared norm of the image-text mean separation, revealing that the mean-separation component is approximately rank-one. A decomposition of the similarity score then identifies three t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית