כתבה
arXiv cs.CL ·
תכנון ויזואלי במצב סמוי
Reason Through the Latent! Making Latent Visual Reasoning Necessary
CVRR הוא שיטה חדשה לתכנון ויזואלי סמוי. היא מאלצת את המודל להשתמש במצב הסמוי לצורך ניבוי. CVRR שומרת על יכולות ויזואליות מוקדמות ומחייבת שימוש בנתונים ויזואליים.
תקציר מקורי באנגליתarXiv:2609.06746v3 Announce Type: replace-cross Abstract: Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply that the model actually relies on that state when producing its answer, especially when alternative image-conditioned paths remain available. We introduce Causal Visual Recurrent Reasoning (CVRR), which preserves pretrained visual competence while making recurrent computation the required image-conditioned path to prediction. CVRR initializes recurrence from the question hidden state after the pretrained vision-language model has incorporated the image, then repeatedly updates this state while re-reading the same fixed
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית