יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

TReVS: תגובה רלוונטית טקסטואלית ושאליות ויזואליות לצמצום טוקן ראייה-לשון

TReVS: Integrating Textual Relevance and Visual Saliency for Efficient Vision-Language Model Token Pruning
TReVS הוא פרקטיקה של צמצום טוקנים של ראייה-לשון שמשלבת רלוונטיות טקסטואלית ושאליות ויזואליות. היא מצמצמת 94.4% מהטוקנים הוויזואליים ושומרת 92.8% מהביצועים של הבסיס הלא מצומצם.
תקציר מקורי באנגליתarXiv:2609.37581v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) excel at visual understanding and reasoning but often incur substantial inference costs due to the large number of visual tokens. Recent visual token pruning methods increasingly follow a two-stage paradigm: they first remove visually redundant tokens after the vision encoder and then discard tokens irrelevant to the textual query within the Large Language Model (LLM). However, since the first stage typically relies solely on vision-encoder saliency, it may prematurely eliminate query-relevant tokens, depriving the subsequent text-guided stage of critical visual evidence. Our empirical analysis shows that incorporating query guidance into first-stage pruning better preserves task-relevant evidence and consisten
קרא במקור המקורי