כתבה
arXiv cs.LG ·
V-CoLA: Vision Token Compression with Linear Attention
תקציר מקורי באנגליתarXiv:2610.11251v1 Announce Type: cross Abstract: Vision-language models (VLMs) have demonstrated impressive capabilities but suffer from substantial computational overhead, as vision tokens dominate the input sequence. This motivates vision token compression as a key direction to alleviate the burden. However, with the emergence of hybrid architectures incorporating linear attention (\eg, Qwen3.5), prior methods designed for softmax attention struggle to generalize. Our analysis reveals that both attention- and similarity-based approaches suffer notable performance degradation, underscoring the urgent need for compression methods tailored to this regime. To this end, we propose \textbf{V-CoLA}, an efficient training-free token compression framework specifically designed for linear attenti
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית