כתבה
arXiv cs.AI ·
חיזוק תהליכי היגיון רב-מודאליים
Reinforcing Multimodal Reasoning via Token-Level Perception-Grounded Advantage Estimation
חוקרים הציגו שיטה חדשה לשיפור יכולות ההיגיון של מודלים רב-מודאליים. השיטה, TPAE, משתמשת במדדים טוקניים כדי לחזק את האותות הרווחים. ניסויים הראו ש-TPAE משפר את היציבות והיעילות של האופטימיזציה.
תקציר מקורי באנגליתarXiv:2609.39168v1 Announce Type: new Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has improved the reasoning capabilities of Multimodal Large Language Models (MLLMs), yet existing frameworks rely on coarse, sequence-level reward signals that lack the fine-grained supervision over the visually-grounded steps within a multimodal reasoning chain. We investigate this gap through the lens of two token-level metrics: visual dependency (i.e. how much a token's prediction relies on the input image features) and predictive entropy. Our empirical analysis reveals two key findings: (1) correct reasoning chains exhibit a markedly sharper entropy reduction as visual grounding intensifies, compared to incorrect ones; (2) pivotal tokens, those whose misprediction triggers reasoning co
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית