כתבה
arXiv cs.AI ·
SLVR: תפיסה נורמלית של רישום חזותי מבוססת על תפיסה אנושית
SLVR: Structured Latent Visual Reasoning via Human-like Reasoning Flows
SLVR מציע תפיסה נורמלית של רישום חזותי, המבוססת על תפיסה אנושית. התפיסה כוללת שלבי תכנון, התקנה, בחירת ראיות ואינטגרציה. SLVR משפר את הביצועים במבחני רישום חזותי, כולל MMVP ו-BLINK Relation.
תקציר מקורי באנגליתarXiv:2610.10563v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) often answer visual reasoning questions by relying on linguistic priors rather than task-relevant visual evidence. Textual chain-of-thought reasoning can partially mitigate this issue by encouraging models to decompose visual questions into intermediate evidence-seeking steps, but generating these steps autoregressively increases inference cost. Latent reasoning avoids explicit rationale generation, but existing approaches provide limited control over what intermediate states encode, making it difficult to impose separate supervision for planning, grounding, and evidence selection. We propose Structured Latent Visual Reasoning (SLVR), a training framework that bridges explicit chain-of-thought and la
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית