יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

MWOP: גזירת פעולות ברוחב מודעת למודלים רב-מודאליים

MWOP: Modality-aware Width-wise Operation Pruning for Efficient MLLMs
MWOP היא שיטה לגזירת פעולות במודלים רב-מודאליים גדולים. היא מאפשרת חיסכון בעלויות חישוב על ידי גזירת פעולות לא הכרחיות. השיטה נבדקה על מודל LLaVA-OneVision-7B והראתה תוצאות טובות.
תקציר מקורי באנגליתarXiv:2610.01434v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) incur substantial inference costs when processing long visual-textual sequences. While existing operation compression methods exploit modality-level redundancy, they largely treat computation within attention heads and shared feed-forward network (FFN) channels as unified units, leaving finer-grained redundancy underexplored. We find that redundancy varies both across modality-interaction paths within the same attention head and across visual and textual executions of the same FFN channel. Based on these findings, we propose Modality-aware Width-wise Operation Pruning (MWOP), which independently prunes visual-to-visual (V2V), text-to-visual (T2V), and text-to-text (T2T) attention paths within each la
קרא במקור המקורי