כתבה
arXiv cs.LG ·
תירוץ דרך הלטנטי
Reason Through the Latent! Making Latent Visual Reasoning Necessary
CVRR הוא שיטה חדשה לתירוץ ויזואלי באמצעות מצב לטנטי. היא מאלצת את המודל להשתמש במידע הוויזואלי הלטנטי לצורך חיזוי. השיטה הוכחה כיעילה במספר בנכות.
תקציר מקורי באנגליתarXiv:2609.06746v2 Announce Type: replace-cross Abstract: Latent visual reasoning aims to perform multimodal reasoning through hidden-state computation rather than explicit textual chains of thought. However, visual information being present in a latent state does not imply that the model actually relies on that state when producing its answer, especially when alternative image-conditioned paths remain available. We introduce Causal Visual Recurrent Reasoning (CVRR), which preserves pretrained visual competence while making recurrent computation the required image-conditioned path to prediction. CVRR initializes recurrence from the question hidden state after the pretrained vision-language model has incorporated the image, then repeatedly updates this state while re-reading the same fixed
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית