כתבה
arXiv cs.LG ·
אין קבוצת שפות אידאלית לאימון הוראה מרוב-לשוני: תצפיות מחקר מומחה
No Optimal Language Set Exists for Multilingual Instruction Tuning: Insights from a Linguistically-Informed Study
לא קיימת קבוצת שפות אידאלית לאימון הוראה מרוב-לשוני. חקירה מומחית חשפה כי קבוצת שפות נבחרת בצורה ספציפית אינה תמידית. כלי: mGPT, mT5-xl, BLOOM.
תקציר מקורי באנגליתarXiv:2410.07809v2 Announce Type: replace-cross Abstract: Multilingual instruction tuning (MIT) is challenged by the curse of multilinguality, data scarcity, and high computational cost. A natural hypothesis is that carefully selecting a linguistically diverse set of languages yields universally better models. We test this systematically by evaluating linguistically-informed selection strategies---based on typological, geographical, semantic, and learned features---against random baselines across three model families (mGPT, mT5-xl, BLOOM) and five multilingual benchmarks. Our key negative finding is that no universal language selection strategy emerges in our fixed-budget setting: performance is strongly task- and model-dependent, and adding more languages beyond a modest threshold trigger
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית