כתבה
arXiv cs.CL ·
דיווחי תכונות תלויות-תצלום במודלי שפה-תמונה
Context-Dependent Affordance Reports in Vision-Language Models
מודלי שפה-תמונה מפיקים תיאורים שונים של עצמים ושימושים תחת פקודות פרסונה שונות.
תקציר מקורי באנגליתarXiv:2603.04419v3 Announce Type: replace Abstract: Vision-language models produce different object and use descriptions under different persona prompts, but low overlap alone does not identify an affordance effect. We audit an earlier seven-prompt study and add matched-question controls. In the historical Qwen pilot, 363 of 3,213 parsed responses contain empty object lists. These affect 2,037 of 9,244 comparisons, with the implementation assigning zero lexical overlap to every affected pair. Conditioning on nonempty reports raises pooled word Jaccard from 0.095 to 0.121 and sentence cosine from 0.415 to 0.511. A previously named chef-specific Tucker factor loses its concentrated loading under missing-cell and complete-nonempty analyses. We withdraw the functional-manifold interpretation a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית