כתבה
arXiv cs.CL ·
Test-Time Training for Modality Order Consistency in Vision-Language Models
תקציר מקורי באנגליתarXiv:2607.20351v1 Announce Type: cross Abstract: We find that vision-language models are sensitive to a specific semantically irrelevant change: the order in which the image and question are presented. Across three models and three benchmarks, image first prompting consistently outperforms question-first prompting, revealing a repeatable modality order failure. We use this gap to design an order-consistent test-time training method. Our method substantially closes the modality-order gap across all evaluated settings. Surprisingly, it also yields consistent improvements in the stronger image-first branch over the baseline, hence bootstrapping both orderings toward mutual consistency. Activation patching localizes the ordering failure to a narrow mid-network region where representations div
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית