יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

MCD: תפיסת קשרים קאוזליים של למידה במצב-מצב במודלי ראייה-שפה גדולים

MCD: Causal Distillation of Multimodal In-Context Learning in Large Vision-Language Models
MCD: תפיסת קשרים קאוזליים של למידה במצב-מצב במודלי ראייה-שפה גדולים. המאמר מציג פרקטיקה של תפיסת קשרים קאוזליים של למידה במצב-מצב במודלי ראייה-שפה גדולים. הפרקטיקה, שנקראת MCD, מעבירה איך שמודל חזק משתמש בראייה ובשפה במהלך למידה במצב-מצב.
תקציר מקורי באנגליתarXiv:2609.39920v1 Announce Type: cross Abstract: Large vision-language models (LVLMs) exhibit strong multimodal in-context learning (ICL) capabilities, yet this ability degrades substantially as model size decreases. Knowledge distillation offers a natural way to bridge this gap, but existing methods primarily align output distributions or hidden representations directly. Such alignment teaches the student what the teacher predicts without revealing which evidence in the complex context causally supports that prediction. Consequently, a student can imitate the teacher's answer while continuing to rely on language priors, prompt structure, or other spurious cues. To address this limitation, we introduce Multimodal Causal Distillation (MCD), a distillation framework that transfers how a str
קרא במקור המקורי