יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

CHRONOBERG: לכידת התפתחות שפה ומודעות זמנית במודלים יסודיים

CHRONOBERG: Capturing Language Evolution and Temporal Awareness in Foundation Models
CHRONOBERG הוא מאגר נתונים מוזמן עם מבנה זמני, המאפשר למודלים יסודיים ללמוד על התפתחות השפה במרוצת הזמן. הוא כולל טקסטים מ-250 שנים, עם עיבודים זמניים שונים. המחקר מראה כי מודלים שאומנו על CHRONOBERG מצליחים פחות לקודד משמעויות זמניות.
תקציר מקורי באנגליתarXiv:2509.22360v2 Announce Type: replace-cross Abstract: Large language models (LLMs) excel at operating at scale by leveraging social media and various data crawled from the web. Whereas existing corpora are diverse, their frequent lack of long-term temporal structure may however limit an LLM's ability to contextualize semantic and normative evolution of language and to capture diachronic variation. To support analysis and training for the latter, we introduce CHRONOBERG, a temporally structured corpus of English book texts spanning 250 years, curated from Project Gutenberg and enriched with a variety of temporal annotations. First, the edited nature of books enables us to quantify lexical semantic change through time-sensitive Valence-Arousal-Dominance (VAD) analysis and to construct hi
קרא במקור המקורי