כתבה
arXiv cs.CL ·
בניית גשרים שפתיים: צירוף נתונים כפילון של כלליות לשכתוב בשפה
Building Multilingual Bridges: Data Mixing as the Pillar of Generalization for In-Language Reasoning
במאמר זה, נציגים פיתוח של L2 reasoning, היכולת של דגם לשכתב בשפה של השאלה. נחקר צירוף נתונים וסדר זמני לשכתוב כללי. נציגים דגם Tiny Aya L2-Thinker בקנה מידה של 3.35B, ומציגים קצב שכתוב של L2 גבוה מ-93% ב-60 שפות על 6 מבחנים.
תקציר מקורי באנגליתarXiv:2609.10445v1 Announce Type: new Abstract: Reasoning language models have made substantial advances on a variety of complex tasks, yet their capabilities remain overwhelmingly English-centric: models primarily reason in English regardless of the language they are prompted in. This is inaccessible for non-English-speaking users, risks losing the intent of the original question, and forgoes knowledge more readily expressed in the target language. In this work, we advance L2 reasoning, the ability of a model to reason consistently in the language of the user's prompt, thus building an in-language bridge between the prompt and the answer. We approach this problem from a data-centric angle, investigating how to optimize data composition and scheduling in SFT for reasoning generalization. B
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית