יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

VETO: אופטימיזציה יעילה של טוקנים למודלים ויז'ואליים-לשוניים

VETO: Video Efficient Token Optimization for Vision Language Models
VETO הוא כלי אופטימיזציה למודלים ויז'ואליים-לשוניים. הוא מקטין את עלות העיבוד של וידאו ארוכים על ידי דחיסה כפולה: דחיסה תוך-מסגרת ודחיסה בין-מסגרתית. VETO משפר את מהירות העיבוד עד 45% תוך שמירה על דיוק.
תקציר מקורי באנגליתarXiv:2610.01785v1 Announce Type: cross Abstract: Processing long videos with Vision-Language Models (VLMs) is bottlenecked by the quadratic cost of visual tokens, making long-form inference prohibitively expensive. While single-axis compression methods mitigate this, they hit a hard efficiency floor because they treat spatial and temporal redundancy independently. We present VETO (Video Efficient Token Optimization for Vision-Language Models), a training-optional plug-in that eliminates this bottleneck through dual-axis compression: (i) an intra-frame compressor that merges semantically similar tokens within each frame via optimal-transport inspired matching, and (ii) an inter-frame compressor that identifies and merges temporally redundant frames. The key design insight is hierarchical o
קרא במקור המקורי