יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

חשוב מחדש: פין-טונינג וקידוד חדש של דגמי Vision-Language

Rethinking Fine-Tuning: Unlocking Hidden Capabilities in Vision-Language Models
מאמר חדש חושף טכניקה חדשה לפין-טונינג של דגמי Vision-Language, שמציעה קידוד יעיל יותר ושיפור בביצועים. הטכניקה, שנקראת Mask Fine-Tuning, משתמשת במסכות לבחירת מידע שנשלח לדגם, ומאפשרת לדגם ללמוד ולשפר את יכולותיו באופן יעיל יותר. המאמר כולל ניסויים שונים שמדגימים את יעילותה של הטכניקה.
תקציר מקורי באנגליתarXiv:2512.23073v2 Announce Type: replace Abstract: Fine-tuning has become the dominant paradigm for adapting Vision-Language Models (VLMs), yet most approaches rely on explicit weight updates that introduce a fundamental trade-off. Full Fine-Tuning (FFT) may perturb pretrained representations due to cross-modal gradient interference, whereas Parameter-Efficient Fine-Tuning (PEFT) methods rely on additive modules, such as low-rank adapters, which may limit adaptation capacity. In this paper, we rethink VLM adaptation from a structural selection framework that adapts VLMs without modifying backbone weights, and we propose Mask Fine-Tuning (MFT). MFT learns masks that selectively route information through existing pretrained connections, dynamically uncovering subnetworks that better align p
קרא במקור המקורי