כתבה
arXiv cs.AI ·
מודל שפה מ-1913: הכשרה על טקסט היסטורי
A Language Model from 1913: Pretraining on Historical Text
הכשרה של מודל שפה על טקסט היסטורי (עד 1913) מציגה תוצאות סבירות בהבנת שפה.
תקציר מקורי באנגליתarXiv:2606.02991v2 Announce Type: replace-cross Abstract: While modern language models increasingly rely on ever-larger web corpora, we show that pretraining on historical text (e.g., pre-1913 text) in a data-constrained setting can produce a temporally grounded language model that still shows reasonable performance on language understanding. However, developing History LMs requires addressing challenges in data quality, preventing temporal leakage in post-training, and constructing temporally aligned evaluations. We address these challenges and pretrain TypewriterLM, a 7.24B-parameter model with a 1913 knowledge cutoff. We construct TypewriterCorpus, a 54B-token historical corpus with extensive temporal filtering, propose lexically grounded instruction tuning that constrains all responses
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית