כתבה
arXiv cs.CL ·
TestHallVQA: חקירת רישום-שפה גדול-מבט-מודלים תחת מצבי תיעוד-שוליים מבחני מדע
TestHallVQA: Exploring LVLMs' Document-Level Reasoning under Redundant Contexts from Scientific Exams
במאמר זה, נוצרה תשתית חדשה למודלי רישום-שפה גדול-מבט, שמטרתה לבדוק את יכולתם להתמודד עם רישומים-תיעודיים גדולים תחת מצבי תיעוד-שוליים. התשתית, הקרויה TestHallVQA, כוללת תרגילים שונים שמטרתם לבדוק את יכולת המודלים להתמודד עם רישומים-תיעודיים גדולים תחת מצבי תיעוד-שוליים. התשתית כוללת גם ניגודים שונים, כגון רישומים-תיעודיים גדולים, רישומים-תיעודיים קטנים, וכדומה.
תקציר מקורי באנגליתarXiv:2609.13158v1 Announce Type: new Abstract: Large Vision--Language Models (LVLMs) are increasingly expected to perform visual question answering (VQA) over planar media. However, existing planar VQA benchmarks typically emphasize isolated challenges: some emphasize long-document understanding with limited reasoning depth, while others require complex visual reasoning but remain restricted to single-page, noise-free settings. Moreover, through theoretical analysis, we identify the impact of irrelevant visual tokens, which leads to measurable performance degradation but has received little attention with respect to systematic quantification. To address these limitations, we introduce TestHallVQA, a multi-image VQA benchmark that simultaneously embodies document-level scale and the diffic
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית