יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

Wieszcz-XIX: קורפוס פולני היסטורי

Wieszcz-XIX: A 3.1-Billion-Word Corpus of Pre-1918 Polish and Temporally Bounded Language Models Trained From Scratch
Wieszcz-XIX הוא קורפוס פולני היסטורי של 3.1 מיליארד מילים. הוא כולל 294,369 מסמכים מ-1800 עד 1918. החוקרים אימנו מודלים מסוג decoder-only על הקורפוס.
תקציר מקורי באנגליתarXiv:2610.10592v1 Announce Type: new Abstract: Historical Polish is well documented as a language but annotated in machine-readable form only to about a million words for the period this paper covers; the rest sits behind optical character recognition of variable quality. We present Wieszcz-XIX, a corpus of 6.75 billion tokens (about 3.1 billion words) in 294,369 documents, most of them periodical issues, of Polish published from 1800 to 1918, assembled from Wolne Lektury and the Internet Archive by a pipeline that filters, deduplicates, audits for post-1918 leakage and splits at the document level. It is over three orders of magnitude larger than the annotated corpus of the same period, and we quantify its defects: recognition corruption against a false-positive floor, near-identical dup
קרא במקור המקורי