כתבה
arXiv cs.CL ·
פיתוח סקאלה של פריתוראי מולטימודלי נטיבי מאפס
Scaling Native Multimodal Pre-Training From Scratch
במאמר זה, נחקרים תכונות הסקאלה של פריתוראי מולטימודלי נטיבי. נמצא כי גודל המודל וכמות הסמנים נקבעים על פי חוק חישוב נכון. נראה כי פריתוראי מולטימודלי נטיבי יכול להשיג יכולות טובות יותר בתחומי תפיסה חזותית ולמידה חזותית-טקסטואלית.
תקציר מקורי באנגליתarXiv:2607.22043v2 Announce Type: replace Abstract: Although large language models (LLMs) exhibit remarkable reasoning capabilities, their reliance on text-only pre-training restricts the perception of the multimodal physical world. Native multimodal pre-training avoids this limitation by training models from scratch on multimodal inputs, thereby achieving deep cross-modal integration and mitigating optimization asymmetries inherent to traditional late-fusion architectures. Despite these advantages, the scaling properties of this paradigm remain incompletely characterized. To address this gap, we investigate the optimal model size and token count for training a Transformer-based vision-language model under a fixed computational budget. Our study demonstrates that minimal objective loss adh
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית