יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מעבר מתיקון לקיצוץ בחישוב ויזואלי

From Patching to Pruning Visual Computation in Vision Language Models
P2P הוא כלי לקיצוץ חישוב ויזואלי במודלים של שפה וראייה. הוא מאפשר לקצר חישובים ללא הפרעה לדיוק. נבדק על מודלים Qwen2.5-VL ו-LLaVA.
תקציר מקורי באנגליתarXiv:2610.03389v1 Announce Type: cross Abstract: Vision language models (VLMs) incur substantial inference cost because every visual token is processed by the attention and MLP projections of every decoder layer, even when token-specific visual computation is unnecessary at many depths. We introduce Patch-to-Prune (P2P), inspired by Mechanistic Interpretability, a training-free framework that converts activation patching from a diagnostic tool into an inference-time computation bypass. P2P performs validation-guided forward and backward layer sweeps to identify decoder regions whose visual-token projection outputs can be replaced by fixed neutral proxy activation vectors within a user-specified accuracy tolerance. Unlike conventional token-pruning methods, P2P preserves the sequence lengt
קרא במקור המקורי