כתבה
arXiv cs.AI ·
LuxInstruct: A Cross-Lingual Instruction Tuning Dataset For Luxembourgish
תקציר מקורי באנגליתarXiv:2510.07074v2 Announce Type: replace-cross Abstract: Instruction tuning has become a key technique for enhancing the performance of large language models, enabling them to better follow human prompts. However, low-resource languages such as Luxembourgish face severe limitations due to the lack of high-quality instruction datasets. Traditional reliance on machine translation often introduces semantic misalignment and cultural inaccuracies. In this work, we address these challenges by creating a cross-lingual instruction tuning dataset for Luxembourgish, without resorting to machine-generated translations into it. Instead, by leveraging aligned data from English, French, and German, we build a high-quality dataset that preserves linguistic and cultural nuances. We provide evidence that
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית