כתבה
arXiv cs.CL ·
CiteVQA: תקן לאטריבוציה של ראיות לביטחון בתבנית תיעודית
CiteVQA: Benchmarking Evidence Attribution for Trustworthy Document Intelligence
CiteVQA הוא תקן שמדורג את האטריבוציה של ראיות לביטחון בתבנית תיעודית. התקן כולל 1,897 שאלות ו-711 תיעודים, ומדורג את המודלים על פי דיוקם באטריבוציה של ראיות.
תקציר מקורי באנגליתarXiv:2605.12882v2 Announce Type: replace Abstract: Multimodal Large Language Models (MLLMs) have significantly advanced document understanding, yet current Doc-VQA evaluations score only the final answer and leave the supporting evidence unchecked. This answer-only approach masks a critical failure mode: a model can land on the correct answer while grounding it in the wrong passage---a critical risk in high-stakes domains like law, finance, and medicine, where every conclusion must be traceable to a specific source region. To address this, we introduce CiteVQA, a benchmark that requires models to return \textit{element-level} bounding-box citations alongside each answer, evaluating both jointly. CiteVQA comprises 1,897 questions across 711 PDFs spanning seven domains and two languages, av
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית