כתבה
arXiv cs.CL ·
Attend to Evidence: Evidence-Anchored Spatial Attention Supervision for Multimodal RLVR
תקציר מקורי באנגליתarXiv:2605.30912v2 Announce Type: replace-cross Abstract: Reinforcement learning with verifiable rewards (RLVR) improves vision-language models (VLMs) by optimizing outcome rewards derived from final answers. However, such outcome-only rewards do not tell the model which image regions justify an answer. For questions that require visual grounding, these rewards cannot distinguish responses supported by relevant visual evidence from those produced by language-prior shortcuts or lucky guesses. We introduce EASE (Evidence-Anchored Spatial Attention), which augments multimodal RLVR with visual-evidence process supervision. EASE converts annotated evidence regions into a smoothed visual-token target and uses it to guide response-to-image attention during RL training, but only on high-reward tra
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית