כתבה
arXiv cs.AI ·
האם עדיין ניתן לעקוב אחר אותות L1?
Can We Still Trace L1 Signals? Investigating the Resilience of Native Language Signals in the LLM Era
חקר השפעת תקופת LLM על אותות שפה ילידית. ניתוח ביצועי זיהוי שפה ילידית בהשוואה לתקופות קודמות. התוצאות מראות ירידה עקבית בביצועים.
תקציר מקורי באנגליתarXiv:2604.08568v3 Announce Type: replace-cross Abstract: The widespread use of LLM-based writing assistance has raised an interesting question about the homogenization of English. As LLMs tend to revise texts toward mainstream English conventions reflected in their training data, the subtle fingerprints that reflect an author's native language (L1) may be gradually disappearing. This study investigates this phenomenon by analyzing native language identification (NLI) performance on academic abstracts. To this end, we construct two NLI datasets of academic abstracts extracted from arXiv and the ACL Anthology that covers eight native language groups across three time periods: pre-neural network (NN), pre-LLM, and post-LLM. We then evaluate NLI performance for each era using NLI classifiers
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית