יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

אי-תואמות חזותיות אינן תואמות חזרה: תכנון זיכוי נמרץ ללמידת רפלקסיה מודאלית

Visual sensitivity is not claim retractability: persistence-aware credit assignment for multimodal reinforcement learning
אפליקציה נמרצת של זיכוי תוך תכנון זיכוי ללמידת רפלקסיה מודאלית. המחקר עוסק בפיתוח תכנון זיכוי שיאפשר למודלים להפצים תוך תכנון זיכוי תוך תכנון זיכוי. המחקר נעשה על ידי Qwen2.5-VL-7B.
תקציר מקורי באנגליתarXiv:2609.36572v2 Announce Type: replace Abstract: Reinforcement Learning with Verifiable Rewards (RLVR) has been extended to Large Vision-Language Models (LVLMs), and perception-aware methods further encourage policies to rely on visual evidence. Yet relying on the image does not guarantee that visual claims are supported by it. Before RL training, 27.81% of the correctly answered responses of Qwen2.5-VL-7B on four multimodal reasoning benchmarks contain at least one direct visual claim that the image does not support. Since outcome-level RL rewards each response as a whole, these claims inherit the positive credit of the correct answer. We introduce a fixed-rollout counterfactual diagnostic that re-scores the same response under an intervened image to separate Evidence-Function Sensitiv
קרא במקור המקורי