כתבה
arXiv cs.CL ·
DocOCR-Eval: A Correction-Based Framework for OCR Tool Selection Without Ground Truth
תקציר מקורי באנגליתarXiv:2607.16203v1 Announce Type: cross Abstract: Document parsing is a foundational step for document understanding tasks such as visual question answering and key information extraction, as it transforms unstructured scanned images into structured representations by extracting textual, visual, and layout information. While numerous Optical Character Recognition (OCR) engines and multimodal large language models (MLLMs) have been developed for this purpose, selecting an appropriate document parsing solution for a given document collection remains challenging, particularly in label-scarce settings. In this work, we conduct a systematic evaluation of text recognition performance across a diverse set of OCR engines and state-of-the-art MLLMs on multiple scanned document benchmarks spanning d
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית