יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

שינויים קצרים, מודלים מורכבים: איך ואיזה תקריות ואי-הספציפיקציה משפיעים על LLMs

Compact Language, Complex Model Shifts: How and Where Ambiguity and Underspecification Affect LLMs
מודלי LLMs נפגעים ביכולת הייצור שלהם עקב תקריות לשוניות ואי-הספציפיקציה. המחקר חוקר את השפעתם של תקריות אלה על המודלים.
תקציר מקורי באנגליתarXiv:2609.39572v1 Announce Type: new Abstract: We analyze how lexical ambiguity and underspecification affect language model training. We create artificial homonyms and artificial hypernyms as pseudowords and analyze the generative performance of language models as they are trained with increasing amounts of these ambiguous or underspecified pseudoword types. We further analyze whether the models disambiguate ambiguous or underspecified statements and provide a first mechanistic account of how ambiguity and disambiguation are represented internally. Our main results show that both ambiguity and underspecification increase model performance in ways that scale with their influence on the language's type-token ratio. However, the accuracy of generating sequences containing ambiguous words or
קרא במקור המקורי