יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

שיפור ייצוג ויזואלי ב-VLM

Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis
חוקרים מציעים שיטה חדשה לשיפור ייצוג ויזואלי במודלים רב-מודאליים (VLM) כמו CLIP. השיטה משתמשת בניתוח שונות קנונית (KCCA) כדי לשפר את הייצוג הוויזואלי.
תקציר מקורי באנגליתarXiv:2610.02718v1 Announce Type: cross Abstract: Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In th
קרא במקור המקורי