יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

חיזוק מודלים ראייה-שפה באמצעות תגמולי ראיות

Making Every Tool Call Count: Necessary Tool-Evidence Path Rewards for Agentic Vision-Language Models
מודלים ראייה-שפה מתקדמים יכולים לענות על שאלות מורכבות, אך לעיתים קרובות הם זקוקים לראיות חיצוניות. NTEP הוא מנגנון תגמול חדש שמעודד מודלים להשתמש בכלים כמו חיפוש תמונות וחיפוש טקסט בצורה יעילה.
תקציר מקורי באנגליתarXiv:2609.03493v1 Announce Type: new Abstract: Modern vision-language models (VLMs) can directly answer many image-grounded questions, yet they often struggle with complex queries requiring fine-grained visual details or external knowledge. To acquire this missing evidence, agentic VLMs invoke tools such as image cropping, image search, and text search. However, existing training paradigms primarily evaluate tool-use based on final answer correctness, leaving evidence acquisition and utilization insufficiently supervised. This leads to two critical shortcomings: (i) models frequently issue redundant or off-target tool calls that fail to gather necessary evidence, and (ii) even when appropriate tools are called, models often fail to extract the necessary information from the resulting obse
קרא במקור המקורי