כתבה
arXiv cs.CL ·
LittleLearner: דגם שפה תחת חשיפה ידע פדגוגית
LittleLearner: Language Models Under Pedagogically Controlled Knowledge Exposure
דגמי שפה שהוכשרו על קורפוס נרכז של חומרי לימוד לבתי ספר יסודיים בארצות הברית.
תקציר מקורי באנגליתarXiv:2608.13545v2 Announce Type: replace Abstract: Modern language models are trained on heterogeneous web-scale text corpora. Consequently, studying knowledge and skill acquisition is difficult, as prior exposure to related content is hard to characterize. To address this challenge, we introduce LittleCurriculum, a curated 88B-token pretraining corpus tailored to U.S. elementary school material, explicitly excluding concepts, facts, and vocabulary taught above Grade 5. Training a 5B-parameter LLM from scratch on LittleCurriculum yields LittleLearner, a model with sufficient language competence for open-ended evaluation, yet with clear knowledge and capability boundaries mapped to interpretable curriculum guidelines. We release LittleCurriculum and LittleLearner as a developmentally restr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית