כתבה
arXiv cs.CL ·
Imagine3D-LLM: הורידו את ה-MLLM ללמוד להתבונן בתמונות 3D
Imagine3D-LLM: Teaching MLLMs to Imagine 3D Scenes Before Answering
המחקר Imagine3D-LLM מציע שיטה חדשה ללמד MLLMs להתבונן בתמונות 3D. השיטה משתמשת בטכניקה של Gaussian Splatting ומספקת תיאור 3D קצר ומקיף של התמונה. המחקר מציע שיטה חדשה ללמד MLLMs להתבונן בתמונות 3D, ומציע תיאור 3D קצר ומקיף של התמונה.
תקציר מקורי באנגליתarXiv:2609.38177v1 Announce Type: cross Abstract: Reasoning about the 3D world from multi-view images remains a fundamental challenge for Multimodal Large Language Models (MLLMs). While modern MLLMs handle single-image inputs effectively, they struggle to integrate evidence across viewpoints into a coherent 3D understanding. A growing body of work attempts to close this gap by injecting 3D awareness into MLLMs, either by boosting fine-grained pixel-level cross-view correspondence or by fusing features from 3D geometry foundation models, yet a substantial gap to human reasoning persists. In this work, we revisit human spatial reasoning, which suggests that rather than relying on fine-grained geometry cues, humans roughly identify common objects across views, infer the relative geometry betw
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית