כתבה
arXiv cs.CL ·
תפקיד הנתונים הסינתטיים ב-OCR תאילנדי
How Far Can Synthetic Data Take Thai OCR?
חוקרים בדקו את היכולת של נתונים סינתטיים לשפר את ביצועי OCR בשפה התאילנדית. הם פיתחו מודל בשם Wayu-Paxa-OCR-Zero, שמשתמש ב-45,723 עמודים סינתטיים ומצליח להפחית את שגיאות האותיות ב-74% בהשוואה למודל הבסיס. המחקר מראה כי נתונים סינתטיים יכולים להיות תחרותיים עם נתונים אמיתיים.
תקציר מקורי באנגליתarXiv:2609.03595v1 Announce Type: new Abstract: We investigate what makes synthetic OCR supervision transfer to real Thai documents and use the resulting insights to build Wayu-Paxa-OCR-Zero, a Thai OCR model adapted without OCR labels from real Thai document pages. Synthetic data provide exact labels at scale, but "realism" conflates source domain, page context, typography, spatial structure, and glyph variation. We disentangle these factors with a controlled document-reconstruction pipeline and evaluate each variant under page- and crop-level training on printed and handwritten Thai documents. Non-text context has little consistent effect, whereas typeface diversity, two-dimensional structure, and real handwriting glyphs improve transfer; moreover, source-domain matching depends on train
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית