כתבה
arXiv cs.LG ·
Seeing Is Not Addressing: Auditing Linguistic Access to Frozen Visual Geometry
תקציר מקורי באנגליתarXiv:2609.37230v1 Announce Type: cross Abstract: Visual distinctions are often finer than those reflected in linguistic conceptualization. Vision-language models exhibit a similar asymmetry: a distinction can remain discriminable in frozen image geometry while being weakly addressable through the native text interface. We study this gap by separating visual discriminability from linguistic addressability in text-to-image retrieval. Using FactorAtlas, a fully crossed testbed of 23,040 images spanning shape, hue, pattern, and nuisance variation, we compare both readouts on held-out images of the same distinctions. We then derive image-side contrasts that separate each value from its alternatives for matched visual grounding, and test whether this reduces the native-text access gap across fa
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית