כתבה
arXiv cs.CL ·
Fusion Anything: דגם כללי למיזוג מודלים מרוב-תכנים
Fusion Anything: A Generalized Multimodal Foundation Model
דגם כללי למיזוג מודלים מרוב-תכנים, המסוגל לעבוד עם תכנים שונים ומשימות שונות. הדגם משתמש בנתונים סינתטיים גדולי-ממדים כדי ללמוד קורלציות מודלי-תכנים. הדגם הזה יכול להיות יעיל יותר מדגמים מיוחדים למשימות מסוימות.
תקציר מקורי באנגליתarXiv:2609.22107v2 Announce Type: replace-cross Abstract: Making prediction with multimodal data is widely used in diverse scenarios. Existing multimodal fusion models, once deployed, can only handle predefined modalities (e.g., vision, text and audio) and single task, making it difficult to quickly adapt to new downstream applications. Therefore, a natural yet aggressive question arises - whether there exists a general multimodal fusion model that can be applied to arbitrary modality combinations and arbitrary prediction tasks. We argue that a unified multimodal fusion model should not depend on specific modalities and should instead encode transferable patterns of multimodal correlation. To this end, we propose a simple and effective learning paradigm based on training on large-scale syn
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית