כתבה
arXiv cs.CL ·
סידור תזונה לגוגל-בי: קוד-שיחת קורסים גורמים להתכנסות חצי-לשונית
Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence
אפשר להכניס התאמה חצי-לשונית למודלי שפה על ידי אימון על טקסט שנקוד-שיחת. נמצא כי אימון על נתונים שנקוד-שיחת גורם להתאמה של תיאוריות פרלליות, ושהתאמה זו נשארת גם לאחר אימון על טקסט יחיד-לשוני. נמצא כי קורסים של קוד-שיחת, התקדמות מרמת מיל-שיחת לרמת משפט-שיחת, גורמים לביצועים טובים יותר של מודלים שאומנו על נתונים שנקוד-שיחת, מאשר של מודלים שאומנו על נתונים שלא נקוד-שיחת.
תקציר מקורי באנגליתarXiv:2609.30535v1 Announce Type: new Abstract: Children in multilingual communities often code-switch, using multiple languages in a single utterance. Can we induce cross-lingual alignment in language models by training on code-switched text? We pretrain small decoder-only transformers on two 100M-word multilingual corpora: a base corpus formed by mixing the English, Dutch, and Chinese BabyBabelLM datasets, and a corpus generated from it by inserting word- and sentence-level code-switching using an LLM. We find that training on code-switched data aligns the representations of parallel text, particularly across different scripts, and that this alignment persists through training on monolingual documents. Under a learning curriculum that progresses from word-level code-switching, to sentenc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית