כתבה
arXiv cs.LG ·
תגמול זהה, מיומנויות שונות: כאשר RL רב-מודאלי לומד להסתכל
Same Reward, Different Skills: When Multimodal RL Learns to Look
RL רב-מודאלי משפר את התוצאות בבדיקות ראייה-שפה, אפילו כאשר המודל אינו מקבל מידע חזותי בשלב האימון. המחקר מראה כי הוספת תמונות בשלב הבדיקה משפרת את התוצאות, אך הדבר דורש מיומנות חדשה מהמודל.
תקציר מקורי באנגליתarXiv:2610.01908v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) improves vision-language benchmark scores even without visual information during training. With images at test, blind-trained models recover roughly half of the real-image gain at 3B and nearly four fifths at 7B. Prolonged real-image training can erode grounding while benchmark gains persist. Both findings expose the same gap: an image in the prompt is not an image in the learning signal. Our design rule, visual resolvability, asks that visual evidence be necessary for a correct answer and that the task remain learnable. We test it on counterfactual coordinate scenes in which the question stays fixed and the target is never named, so a correct answer requires finding the target in the imag
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית