כתבה
arXiv cs.LG ·
Reason Through the Latent! Making Latent Visual Reasoning Necessary
תקציר מקורי באנגליתarXiv:2609.06746v1 Announce Type: cross Abstract: Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply that the model actually relies on that state when producing its answer, especially when alternative image-conditioned paths remain available. We introduce \textbf{C}ausal \textbf{V}isual \textbf{R}ecurrent \textbf{R}easoning (CVRR), which preserves pretrained visual competence while making recurrent computation the required image-conditioned path to prediction. CVRR initializes recurrence from the question hidden state after the pretrained vision-language model has incorporated the image, then repeatedly updates this state whil
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית