כתבה
arXiv cs.CL ·
Evidence Attribution in Visual Document Understanding without Coordinates or Region Labels
תקציר מקורי באנגליתarXiv:2607.24651v1 Announce Type: cross Abstract: Reliable visual document understanding requires a model to attribute each answer to the evidence regions that support it. Recent benchmarks and systems express this step through a coordinate interface: the model outputs the coordinates of bounding boxes that mark the evidence regions in the document. Under this interface, vision-language models often fail to identify the right regions even when the answer is correct, a failure known as Attribution Hallucination. We present a study that investigates whether this failure is partially limited by what the model can express through coordinates. On a verified bilingual CiteVQA subset, we compare the coordinate interface with a language interface in which the model outputs only text, quoting its e
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית