כתבה
arXiv cs.AI ·
Decoupled Visual Processing: Efficient Multimodal Adaptation via Modality-Specific Transformer Substitution
תקציר מקורי באנגליתarXiv:2607.26596v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have demonstrated remarkable capabilities by integrating visual and textual understanding within a unified transformer architecture. However, fine-tuning all parameters of these models for visual instruction tuning is computationally expensive and often unnecessary, as the representation requirements for visual and textual tokens diverge significantly in the deeper layers of the network. In this paper, we propose Decoupled Visual Processing (DVP), an efficient training framework that replaces the upper decoder layers of a pretrained LLM with a lightweight, independently trainable single transformer block dedicated exclusively to visual token processing. Specifically, after shared processing through t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית