יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

כאשר להתאים, וכאשר לנבא: תרשים פאזה ללמידה מודאלית

When to Align, When to Predict: A Phase Diagram for Multimodal Learning
תרשים פאזה שמסווג את המטרות הלמידה המודאלית לארבעה תחומים: Both, CA only, CP only, ו-Neither. המאמר עוסק בהשוואה בין CA ו-CP, ובכך שהם נותנים תוצאות שונות בתחומים שונים. התרשים נוצר תחת תחום חיפוש ספייקד, ומציג תחומי נישה שונים. המאמר כולל ניסויים על נתונים סינתטיים, תצלומים-תיאורים, ותחומים מדעיים שונים.
תקציר מקורי באנגליתarXiv:2606.11190v3 Announce Type: replace Abstract: Cross-modal alignment (CA) and cross-modal prediction (CP) are the dominant paradigms for multimodal representation learning, yet there is no systematic understanding of when each succeeds and when each fails --- a gap that leaves practitioners, especially in scientific domains with heterogeneous instruments and multiple levels of measurement, unable to diagnose why standard methods underperform the best single modality. We study both objectives under a spiked signal-plus-noise model with structured cross-modal nuisance correlation, the ingredient that breaks the classical recovery guarantees, and derive separation ratios that expose complementary failure modes: alignment whitens each modality and fails when nuisance is strongly correlate
קרא במקור המקורי