יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Reason Through the Latent! Making Latent Visual Reasoning Necessary

מודל המחשבה ויזואלית סמויה CVRR שומר על יכולת ראיית עצמאית מוקדמת, ומחייב חישוב חוזר כדרך התנאית לקבלת תשובה.
תקציר מקורי באנגליתarXiv:2609.06746v2 Announce Type: replace-cross Abstract: Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply that the model actually relies on that state when producing its answer, especially when alternative image-conditioned paths remain available. We introduce Causal Visual Recurrent Reasoning (CVRR), which preserves pretrained visual competence while making recurrent computation the required image-conditioned path to prediction. CVRR initializes recurrence from the question hidden state after the pretrained vision-language model has incorporated the image, then repeatedly updates this state while re-reading the same fixed
קרא במקור המקורי