כתבה
arXiv cs.CL ·
MinerU-Popo: Universal Post-Processing Model for Structured Document Parsing
תקציר מקורי באנגליתarXiv:2605.24973v2 Announce Type: replace-cross Abstract: VLM-based OCR models have become the de facto choice for document parsing, as they can accurately extract page-level elements (e.g., paragraphs within individual pages) together with their bounding boxes and textual content. However, downstream applications such as RAG require coherent document-level information, whereas these models often break cross-page continuity and fail to recover disrupted structures, such as paragraphs and tables truncated by page boundaries. Such relationships are not confined to a single page; instead, they require joint analysis of titles, paragraphs, tables, and images spanning multiple pages. A natural solution is therefore to reuse existing OCR outputs and reconstruct document-level logical structures
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית