כתבה
arXiv cs.LG ·
WOVEN: רקיחת תיאור עולם ויזואלי ל-LMMs
WOVEN: Weaving Visual World Modeling into Multimodal LLMs
LMMs המודעים למודלים סבלו מקשיי תפיסה חללית. WOVEN מציגה מקור ובסיס ניסוי לתיאור תפיסה ויזואלית.
תקציר מקורי באנגליתarXiv:2610.12417v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) struggle with spatial, embodied, physical, and temporal reasoning. We hypothesize that these failures reflect a shared deficit in visual transition reasoning, and test whether this capability can serve as a shared training primitive, one that different models can learn from different supervision sources and reuse across different tasks, with a systematic training recipe. Existing benchmarks document these deficits separately but do not support controlled comparisons across scenes, actions, and reasoning operations. We therefore introduce WOVEN, a training source and benchmark for visual transition reasoning that organizes transition supervision by scene, action, and reasoning type, using diverse, rea
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית