יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

ייצור נתונים סינתטיים לתרגום מכונה

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books
חוקרים פיתחו שיטה לייצור נתונים סינתטיים לתרגום מכונה בשפות עם משאבים מוגבלים. השיטה משתמשת במודלים גדולים לשפה כדי לחלץ כללים דקדוקיים וליצור נתונים סינתטיים. הניסויים הראו שיפור בתוצאות התרגום.
תקציר מקורי באנגליתarXiv:2607.22376v2 Announce Type: replace Abstract: Most endangered languages lack the parallel data required for machine translation, despite the existence of descriptive grammar books. We introduce a pipeline that uses large language models to extract grammatical rules, example sentences, and lexicons from grammar books and generate synthetic parallel corpora for fine-tuning-rather than feeding grammar content into prompts at inference time, as in prior work. Validated on three typologically diverse low-resource languages-Kalamang (Papuan), Tuatschin (Romance), and Mandan (Siouan)-we show that fine-tuning on synthetic data improves over seed-data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, with best-case ChrF++ gains of +8.8, +5.3, and +3.3 respectively. Throug
קרא במקור המקורי