כתבה
arXiv cs.LG ·
שילוב נתונים סינתטיים ואמיתיים ל-OCR היסטורי: מקרה סטודיו של מאנג'ו
Combining Synthetic and Real Data for Low-Resource Historical OCR: A Manchu Case Study
שילוב נתונים סינתטיים ואמיתיים ל-OCR היסטורי: ניתוח קסדה של מאנג'ו. ניתוח קסדה של מאנג'ו.
תקציר מקורי באנגליתarXiv:2609.11495v1 Announce Type: new Abstract: Manchu, now critically endangered, was one of the principal languages of the Qing empire (1636-1912), and its extensive archival record is increasingly digitized but remains difficult to search and analyze at scale. Previous work showed that vision-language models (VLMs) trained only on synthetic Manchu word images can reach 87.4% word accuracy on real Qing manuscripts and prints, leaving a substantial synthetic-to-real gap. This study examines how synthetic and real historical training data should be combined for low-resource OCR. Using 60,000 synthetic and 20,306 real historical word images, we evaluate three pretrained VLMs and a compact convolutional recurrent neural network (CRNN) under four regimes: synthetic-only, real-only, joint synt
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית