כתבה
arXiv cs.LG ·
DIPrune: Task-Aware Token Pruning with Dual Importance for Efficient Multimodal Language Models
תקציר מקורי באנגליתarXiv:2610.08341v1 Announce Type: cross Abstract: Recent training-free pruning approaches for Multimodal Large Language Models (MLLMs) effectively cut computational overhead by exploiting visual redundancy or text-vision attention. However, they frequently suffer from semantic degradation due to their task-agnostic design or unreliable attention estimates. Based on our empirical analysis, we have found that this issue arises because salient tokens in shallow layers persistently suppress emerging semantic ones through numerical inertia, leading to premature discarding of signals crucial for deep reasoning. To address the aforementioned issue, from the task-oriented aspects, we first reformulate training-free pruning as a minimization of the distortion in the final task loss and derive a tra
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית