כתבה
arXiv cs.CL ·
קדימונים מבניים ללמידת שפה יעילה
Structural priors for data-efficient language learning
חוקרים בודקים שיטות להפחתת התלות בנתונים גדולים ומשאבים חישוביים ללמידת שפה. הם מציעים 'קדימונים מבניים' - אימון מודלים על נתונים לא-לשוניים כדי ליצור ידע קודם ללמידת שפה טבעית. התוצאות מראות כי נתונים סימבוליים, כגון מוזיקה ותכנות, יכולים להפחית את ההפסד בלמידת שפה.
תקציר מקורי באנגליתarXiv:2609.11505v1 Announce Type: new Abstract: Efficient language learning requires methods to reduce the reliance on large data and computational resources. We investigate structural transfer: First training models on non-language data to induce useful priors for natural language. This approach is a form of weight initialization for multilingual language modeling. We evaluate transfer via next-token-prediction loss, weight shifts in the model, and downstream linguistic benchmarks. Several symbolic data types - notably music, probabilistic grammars, and cellular automata - yield lower language-modeling loss than random initialization. These gains coincide with smaller weight shifts during subsequent language training, suggesting that structural transfer positions models in a more favorabl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית