יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

אפיון תוכן נרטיבי בנתוני אימון LLM

Characterizing Narrative Content in Web-scale LLM Pretraining Data
חוקרים בדקו את התוכן הנרטיבי בנתוני אימון LLM. הם פיתחו מסגרת לאפיון תוכן נרטיבי ויישמו אותה על 13M קטעי טקסט. התוצאות מראות שאיכות התוכן הנרטיבי משתנה בין מקורות ונושאים שונים.
תקציר מקורי באנגליתarXiv:2606.19468v2 Announce Type: replace Abstract: The narrative composition of web-scale LLM pretraining corpora remains largely unexplored, even though narrative is a fundamental mode of human communication. We present the first fine-grained study of narrative features in Dolma, a 3-trillion-token open pretraining corpus. Drawing on narrative theory, we design a framework spanning three core narrative elements (agency, setting, and events) operationalized as 11 interpretable dimensions. After curating and hand-annotating a diverse set of 400 passages, we create an LLM-labeled dataset of 25K passages, and finally, we finetune and validate NarraBERT, two RoBERTa-based models for fine-grained narrative prediction. We apply NarraBERT to 13M passages, resulting in a new dataset, NarraDolma.
קרא במקור המקורי