כתבה
arXiv cs.AI ·
שיפור ייצוג ויזואלי ב-VLMs
Revisiting Visual Representation Enhancement of VLMs via Kernel Canonical Correlation Analysis
חוקרים מציגים שיטה חדשה לשיפור ייצוג ויזואלי במודלים של שפה וראייה, כגון CLIP. השיטה משתמשת בניתוח שונות קאנוני גרעיני (KCCA) כדי לשפר את הייצוג הוויזואלי.
תקציר מקורי באנגליתarXiv:2610.02718v1 Announce Type: cross Abstract: Vision-language models such as CLIP exhibit strong semantic generalization, but remain limited in fine-grained visual perception. A recent work named KUEA presents a natural remedy by finetuning the image encoder under the supervision of the vision-centric DINOv2 to align their kernel matrices element-wisely, while regularizing the embeddings to remain close to the pretrained visual encoder for preserving image-text semantics in CLIP. However, we show that diminishing the role of the alignment loss to DINOv2 does not necessarily degrade its fine-grained visual performance, suggesting that the kernel-matrix discrepancy may be insufficient for further visual representation enhancement, motivating us to revisit the alignment formulation. In th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית