כתבה
arXiv cs.AI ·
ITO: סינכרון ואיחוד של תצוגות רב-מודליות ואימון-זמן-איחוד לאימון טקסט-תמונה
ITO: Multi-View Alignment and Training-Time Fusion for Image-Text Pretraining
אימון טקסט-תמונה רב-מודלי עם סינכרון ואיחוד של תצוגות רב-מודליות ואימון-זמן-איחוד. הפרקטיקה ITO מציעה פתרון לבעיית האימון של תצוגות רב-מודליות. הפרקטיקה ITO מציעה פתרון לבעיית האימון של תצוגות רב-מודליות. הפרקטיקה ITO מציעה פתרון לבעיית האימון של תצוגות רב-מודליות.
תקציר מקורי באנגליתarXiv:2603.02767v4 Announce Type: replace-cross Abstract: Image--text contrastive pretraining has become a dominant paradigm for visual representation learning, yet existing methods often yield representations that remain partially organized by modality rather than by semantics. We propose ITO, a framework addressing this limitation through two complementary mechanisms with distinct roles. Multimodal multiple alignment enriches supervision by constructing diverse cross-modal correspondences from multi-view image augmentations, providing the primary source of discriminative gain. A lightweight training-time multimodal fusion module then acts as a training-time regularizer, encouraging the encoders to produce features that are compatible under fusion. Crucially, the fusion module is discarde
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית