יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חשוב מחדש את האובייקטיב ההשוואתי ב-CLIP לאחר הכשרה: פרקטיקה משלימה עם מפיץ טקסט קפוא

Rethinking Contrastive Loss in CLIP Post-training: A Complementary Framework with Frozen Text Encoder
אנו חושבים מחדש את האובייקטיב ההשוואתי ב-CLIP לאחר הכשרה, ומציעים פרקטיקה משלימה שמשפרת את המודל. הפרקטיקה, שנקראת ComCLIP, משתמשת באובייקטיב ההשוואתי המתוקן ובאובייקטיב MSE, ומציעה תוצאות טובות יותר מאשר המודל המקורי. ComCLIP נבחן באמצעות 8 בדיקות, והתוצאות היו טובות. הפרקטיקה נמצאה להיות יעילה ומשפרת את המודל.
תקציר מקורי באנגליתarXiv:2610.11374v1 Announce Type: cross Abstract: CLIP serves as a foundational vision-language model and the de facto vision encoder for downstream VLMs such as LLaVA. Post-training offers a lightweight route to refine CLIP, but recent work argues that the standard contrastive loss is unsuitable for post-training due to catastrophic forgetting under small batches, motivating designs that abandon the contrastive objective in favor of distillation. We revisit this premise and find that, for the InfoNCE objective, the reported forgetting is driven primarily not by insufficient negatives but by an inappropriate magnitude of the contrastive temperature $\tau$: with $\tau$ set sufficiently small, contrastive post-training improves rather than degrades the pretrained CLIP, which we explain throu
קרא במקור המקורי