כתבה
arXiv cs.LG ·
Visual Branch היא הדבר שאתה צריך לצאת לפעולה עבור CLIP-based Class-Incremental Learning
Visual Branch is What You Need for CLIP-based Class-Incremental Learning
למדדי CLIP-based Class-Incremental Learning יש צורך במתודולוגיה ראשונה שלא תשתמש בברנך טקסטואלי ותבנה את המחלקה הכל-כללית כולה במרחב ה-Visual.
תקציר מקורי באנגליתarXiv:2609.37888v2 Announce Type: cross Abstract: Class-Incremental Learning (CIL) requires models to recognize new classes over time without forgetting previously learned ones. With the rise of vision-language pre-training, CLIP has become a strong foundation for CIL. A common design in CLIP-based CIL is to construct textual classifier weights by encoding class-name templates with the CLIP text encoder, and then classify visual features by image-text cosine similarity. This design is appealing: since CLIP aligns images and text in a shared embedding space, textual weights appear to provide an off-the-shelf classifier for incremental classes. However, we show that this seemingly natural design is not always beneficial, as a modality gap can still separate the two modalities and make textua
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית