כתבה
arXiv cs.CL ·
Gaokerena: A Small Persian Medical Language Model Family
תקציר מקורי באנגליתarXiv:2608.00932v4 Announce Type: replace Abstract: The integration of artificial intelligence into medical question-answering systems has advanced rapidly; however, research remains predominantly focused on English, leaving low-resource languages like Persian significantly underserved. To address this gap, this paper introduces Gaokerena, a novel family of compact Persian medical language models optimized for deployment on consumer-grade hardware. As a foundational step toward localized digital healthcare, we first present Gaokerena-V, developed by training a baseline model on a strategically selected subset of a newly curated 90-million-token Persian medical corpus (approximately 54 million tokens) together with 20,000 expert-vetted physician Q&A pairs (approximately 3 million tokens), f
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית