יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

CHRONOBERG: תיעוד התפתחות שפה והווה-מודעות זמנית במודלי יסוד

CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models
קורפוס Chronoberg תיעוד התפתחות שפה והווה-מודעות זמנית במודלי יסוד. Chronoberg זמין ב-HuggingFace ובגיטהאב.
תקציר מקורי באנגליתarXiv:2509.22360v2 Announce Type: replace Abstract: Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web. Whereas existing corpora are diverse, their frequent lack of long-term temporal structure may however limit an LLM's ability to contextualize semantic and normative evolution of language and to capture diachronic variation. To support analysis and training for the latter, we introduce CHRONOBERG, a temporally structured corpus of English book texts spanning 250 years, curated from Project Gutenberg and enriched with a variety of temporal annotations. First, the edited nature of books enables us to quantify lexical semantic change through time-sensitive Valence-Arousal-Dominance (VAD) analysis and to construct historic
קרא במקור המקורי