כתבה
arXiv cs.CL ·
WOVEN: שילוב דגמי עולם חזותיים ב-LLM רב-מודאליים
WOVEN: Weaving Visual World Modeling into Multimodal LLMs
WOVEN הוא מקור אימון ובנך' לתהליכי מעבר חזותיים, המאפשר למודלים רב-מודאליים ללמוד יכולות חדשות. המחקר בדק 38 מודלים, כולל GPT-5.4 ו-Qwen3-VL-235B-A22B, ומצא פער משמעותי ביכולות הספציפיות. האימון על WOVEN שיפר את התוצאות ב-22 מתוך 26 בנך'ים חיצוניים.
תקציר מקורי באנגליתarXiv:2610.12417v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, rea
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית