כתבה
arXiv cs.CL ·
Jina-OCR-v1: פארסינג תיעוד יעיל
Jina-OCR-v1: Efficient Document Parsing with Speculative Decoding and Dense Verifiable Rewards
Jina-OCR-v1 הוא מודל פארסינג תיעוד יעיל שמשלב את DeepSeek-OCR. המודל משיג תוצאות טובות על OmniDocBench v1.6 ו-olmOCR-Bench, ומגיע לקצב עמודים הגבוה ביותר בהשוואה.
תקציר מקורי באנגליתarXiv:2609.03181v1 Announce Type: new Abstract: We present Jina-OCR-v1, an end-to-end document parsing model built to serve on low-budget GPUs. It combines the compressed-vision encoder and the 3B mixture-of-experts decoder of DeepSeek-OCR, which activates about 570M parameters per token, with a FastMTP speculative decoding head that shares a single draft block recursively across K=3 prediction steps. Greedy verification makes decoding lossless. Post-training combines instruction alignment, robustness fine-tuning on difficult documents, and GRPO under dense verifiable rewards: deterministic formula, table, and structural checks that award partial credit. The training data mixes cleaned public corpora with targeted synthetic pages. At the default dynamic-resolution setting, Jina-OCR-v1 scor
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית