כתבה
arXiv cs.CL ·
למידה של תצפיות נוספות במהלך הכשרה קדם-אימון: דגלי שפה גדולים לומדים טוב יותר עם תצפיות נוספות
Knowledge Acquisition During Pre-training? Large Language Models Learn Better With Auxiliary Views
דגלי שפה גדולים לומדים טוב יותר כאשר הם נחשפים לתצפיות נוספות במהלך הכשרה קדם-אימון. תצפיות אלו יכולות להיות תרגומים של המידע, והן עשויות לסייע לדגלי השפה ללמוד טוב יותר.
תקציר מקורי באנגליתarXiv:2609.04180v1 Announce Type: new Abstract: Gaps remain in our understanding of how large language models (LLMs) acquire knowledge during pre-training. We posit that auxiliary views, reformulations of knowledge, are causally helpful for learning. We design controlled experiments to isolate this. First, we confirm that repetition is necessary for acquisition and clarify that paraphrasing helps only at smaller batch sizes. Second, holding the token budget fixed, allocating tokens from document repetition to auxiliary views improves learning, counterintuitively, even for factual recall. Third, the effectiveness of auxiliary views is not contingent on the strength of the teacher model that generates them. Fourth, we identify forms of knowledge, contextual and foundational, that aid learnin
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית