יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מפיקים מפיקסלים למבנה: דגמי ראייה-שפה קלים ל-OCR והפיכת JSON מבורך

From Pixels to Structure: Lightweight Vision-Language Models for Document OCR and Structured JSON Extraction
דגמי ראייה-שפה קלים ל-OCR והפיכת JSON מבורך. חידוש: דגמי ראייה-שפה קלים ל-OCR והפיכת JSON מבורך. כולל: LangGraph, GPT-5, Gemini.
תקציר מקורי באנגליתarXiv:2610.11818v1 Announce Type: cross Abstract: While massive, closed-source Vision-Language Models (VLMs) set strong benchmarks for document understanding, their dependence on commercial APIs limits adoption in institutional archives due to data autonomy concerns, recurring costs, and the environmental footprint of hyperscale computing. This is especially acute in heritage digitization, where documents include historical handwriting, domain-specific terminology (e.g., jewelry, prehistory, architecture), and non-standard layouts requiring high-dimensional structured extraction. We present a comparative study of eight open-source lightweight VLMs (up to 7B parameters) for Optical Character Recognition (OCR)-to-structure across three university heritage collections. Given a document image,
קרא במקור המקורי