כתבה
arXiv cs.CL ·
כיול מבוסס אדם להתאמה בין טקסט ארוך לתמונות
Human-Grounded Calibration for Long-Text Image-Text Congruence in Vision-Language Models
חוקרים הציגו שיטה חדשה לכיול התאמה בין טקסט ארוך לתמונות. השיטה, Congruency Score, מאפשרת למודלים לחשב ציון התאמה מדויק יותר. המחקר בדק ארבעה מודלים שונים ומצא שהשיטה החדשה משפרת את היכולת להתאים טקסט לתמונות.
תקציר מקורי באנגליתarXiv:2609.15640v1 Announce Type: cross Abstract: Long-text image--text congruence scoring is increasingly important for vision-language systems that must evaluate whether detailed textual descriptions match visual content. However, raw similarity scores from dual-encoder models are difficult to interpret as calibrated congruence measures, especially under the modality gap between image and text embeddings. This paper proposes Congruency Score (CS), a lightweight calibration layer that maps image--text similarity evidence into a bounded score. Using DOCCI and Urban1k, we evaluate four frozen vision-language backbones and show that observed reductions in post-projection centroid distance do not uniformly improve image--text retrieval performance. Human-grounded evaluations on DOCCI further
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית