כתבה
arXiv cs.AI ·
Learning to Predict Middle-Layer Attention in MLLMs for Visual Token Pruning
תקציר מקורי באנגליתarXiv:2608.06411v2 Announce Type: replace Abstract: Multimodal large language models (MLLMs) achieve strong performance across diverse vision-language tasks, but their efficiency is limited by the cost of processing numerous visual tokens. Visual token pruning can reduce this cost, but requires accurate token importance estimates. Recent studies have demonstrated that text-to-vision attention from middle language model layers can effectively guide visual token pruning, typically using attention from a predefined middle layer to select the visual tokens to retain. Two problems therefore remain. First, our analysis shows that the layer whose attention is most responsive to the question varies substantially across samples, making a fixed layer suboptimal. Second, obtaining attention from the
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית