כתבה
arXiv cs.CL ·
From Zero to Hero: An Open LLM Ecosystem for Armenian
תקציר מקורי באנגליתarXiv:2609.03350v1 Announce Type: cross Abstract: Pretraining data for Armenian, a morphologically rich and low-resource language, is scarce, and no open Armenian LLM has been released with the data and recipe needed to reproduce it. To address this gap, we curate and release two datasets. ArmWeb is an extensively validated corpus of 4.37M Armenian news documents. ArmSTEM is a parallel English-Armenian collection of 373K math and science problems with step-by-step solutions, translated into Armenian and verified through both answer-preserving LLM judgment and human evaluation. Continued pretraining of Gemma-4-E4B on these datasets yields arm-gemma-e4b, which outperforms every existing open Armenian model as well as its unadapted base, and is the first open Armenian LLM with complete traini
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית