כתבה
arXiv cs.LG ·
Harnessing Image Question Dependence for Better VLM Test-time Reinforcement Learning
תקציר מקורי באנגליתarXiv:2609.13296v1 Announce Type: cross Abstract: Test-time reinforcement learning can adapt vision-language models (VLMs) to unlabeled target data, but its effectiveness is fundamentally limited by the reliability of self-generated learning signals. To assess the reliability of consensus-based learning signals, we analyze VLM test-time reinforcement learning across diverse VQA datasets and model sizes, revealing two limitations. First, gains from consensus-based test-time training largely come from answer normalization rather than content correction. Second, many initial VLM responses are incorrect due to the model's limited ability to jointly use the image and the question; consensus rewards derived from these outputs may preserve the resulting grounding errors rather than correct them.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית