כתבה
arXiv cs.CL ·
מגרפים ידע ספציפי לשאלות לתבונה ראייתית יעילה
Question-Specific Knowledge Graphs for Efficient Visual Reasoning
מגרפים ידע ספציפי לשאלות מסייעים לתבונה ראייתית יעילה. המאמר עוסק בפיתוח של VisKG, פלטפורמה של RL שמטרתה לשפר את תבונת הראייה של דגימות תמונה. VisKG משתמש במגרפים של ידע ספציפי לשאלות, ומטרתו לשפר את תבונת הראייה של דגימות תמונה. המאמר עוסק בפיתוח של VisKG, ובבדיקות של VisKG על ידי המחברים.
תקציר מקורי באנגליתarXiv:2609.35942v1 Announce Type: new Abstract: Recent work in visual question answering has shown that vision-language models can exhibit strong reasoning capabilities by translating visual inputs into textual representations. The effectiveness of this translation depends on how well visual details are retained; models need to surface and align both explicit and implicit knowledge sufficient to support reasoning, without introducing spurious assumptions. Existing methods that leverage detailed image captions introduce visual details unrelated to the reasoning task, inflating input token counts and increasing computational cost. To address these challenges, we propose VisKG, a reinforcement learning (RL) framework in which models learn to translate visual content into question-specific kno
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית