כתבה
arXiv cs.AI ·
Dual-Latent Memory Routing for Vision-Language Reasoning
DLMR משפר את התפיסה-שפה ב-MLLMs. המנוע כולל זיכרון חצי-סתום וזיכרון סיבוב, ומאפשר רשומה עצמאית של ראיות ומסקנות. ה-DLMR מוכנס בשלושה שלבים, ומשפר את הביצועים בשני תחומים: כללי ותפיסה.
תקציר מקורי באנגליתarXiv:2609.05539v1 Announce Type: cross Abstract: Multimodal large language models (MLLMs) have recently made strong progress in vision-language reasoning, yet their performance often degrades as generations grow longer. A key factor is that they frequently lose track of earlier visual evidence and intermediate constraints under a monolithic growing context. Inspired by how humans separately recall what they see and what they infer when solving complex tasks, we propose DLMR, a parameter-efficient mechanism that equips MLLMs with Dual Latent Memories: a visual memory that compresses image evidence and a reasoning memory that tracks intermediate conclusions and constraints. A Router then dynamically decides which memory and how much to reuse during inference, preserving visual grounding whi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית