יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מודלי LLM נוטים לסדר מילים משמאלה: ניתוח נרחב של 192 שפות

Left-Branching Transformers Excel at Right-Branching Languages: Data Shapes Word Order Preferences in Language Models
מודלי LLM נוטים לסדר מילים משמאלה בשפות סינתטיות, אך משמאל-ימין בשפות טבעיות. המחקר חושף קשר בין סדר המילים למאגרי הנתונים ואיכותם.
תקציר מקורי באנגליתarXiv:2608.15129v2 Announce Type: replace Abstract: We systematically compare word order preferences in decoder-only language models across 192 artificial languages and typologically diverse natural languages. On artificial languages, models exhibit a left-branching preference that aligns with neither natural language universals nor human word order learning biases. On natural languages, monolingual models show no clear base word order bias at small scales, but as data grows, a preference for right-branching subject-verb-object (SVO) languages emerges while SOV falls behind despite being the most frequent order cross-linguistically. This SVO advantage extends to multilingual models and correlates with language resource level and data quality rather than word order. Thus, the same architect
קרא במקור המקורי