יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מפיצי טקסט גדולים יכולים לפגוע בביצועי CLIP Zero-Shot

Bigger Text Encoders Can Hurt CLIP Zero-Shot Performance
מפיצי טקסט גדולים יכולים לפגוע בביצועי CLIP Zero-Shot. נמצא כי גודל המפיץ הטקסטואלי יכול להשפיע על הביצועים של CLIP, ושהגדלת המפיץ הטקסטואלי יכולה לגרום להתפשטות מוגברת. נמצא גם כי שימוש במכפלה ספציפית לכוחות נגד זיהוי עוזר לשפר את הביצועים.
תקציר מקורי באנגליתarXiv:2609.05730v1 Announce Type: cross Abstract: Contrastive Language-Image Pretraining (CLIP) is a building block of many machine learning applications. Scaling laws have guided resource allocation for large-scale training, yet prior work treats total CLIP model size as a single variable, without exploring how the capacity split between encoders impacts downstream performance. Here, we train multiple CLIP models with different vision and text encoder sizes, revealing that for most vision encoders, there is an optimal text encoder size beyond which zero-shot performance degrades---even as total parameter count increases. Exploiting this behavior yields efficient configurations that match the zero-shot performance of the standard ViT-B/16 architecture with up to 55% fewer parameters. We fu
קרא במקור המקורי