כתבה
arXiv cs.LG ·
Architecture-Dependent Fusion Pathways in MLLMs
תקציר מקורי באנגליתarXiv:2610.03289v1 Announce Type: new Abstract: Multimodal Large Language Models (MLLMs) achieve strong performance across vision-language tasks, yet the internal mechanisms by which visual and textual information are fused across layers remain insufficiently understood. We investigate representative MLLMs from two architectural paradigms: concatenation architectures and native multimodal architectures. We conduct three progressively connected analyses: alignment decoupling identifies which modality changes, attention routing and entropy characterize how cross-modal information is distributed, and intrinsic dimensionality examines how fusion reshapes feature spaces. Separately, we perform causal intervention experiments as a validation of the resulting interpretation. As a supplementary an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית