כתבה
arXiv cs.CL ·
TEVI: עריכה מותנית טקסט של ייצוגים חזותיים
TEVI: Text-Conditioned Editing of Visual Representations via Sparse Autoencoders for Improved Vision-Language Alignment
TEVI הוא כלי לעריכה מותנית טקסט של ייצוגים חזותיים. הוא משתמש בקפטיונים כאות לזיהוי מה לשמור מתוך ייצוגים חזותיים. TEVI משפר את ההתאמה בין ייצוגים חזותיים וטקסט, ומאפשר שיפור בביצועים.
תקציר מקורי באנגליתarXiv:2606.07451v2 Announce Type: replace-cross Abstract: Vision-language models such as CLIP are highly useful for diverse tasks due to their shared image-text embedding space. Despite this, the image and text embeddings are often poorly aligned, affecting downstream performance. Recent work has hypothesized that this can be attributed to an information imbalance: images contain more information than their captions describe. In this work, we propose TEVI, a framework that uses captions as a signal for what to retain from image embeddings. Specifically, we use sparse autoencoders to disentangle image embeddings and train a masking module to selectively reconstruct the embedding based on a given caption. In a controlled setup with synthetic captions, we show that TEVI is effective at preser
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית