יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

IndicTriMix: פיתוח קבצי נתונים ומודלים לזיהוי שפה בטקסט מעורב

IndicTriMix: Developing Language Identification Datasets and Models for Tri-Language Code-Mixing
ניתן לקרוא על פיתוח קבצי נתונים ומודלים לזיהוי שפה בטקסט מעורב. המאמר עוסק בפיתוח מודלים לזיהוי שפה בטקסט מעורב, כולל שמות מודלים/כלים/חברות. המאמר עוסק בפיתוח מודלים לזיהוי שפה בטקסט מעורב, כולל שמות מודלים/כלים/חברות.
תקציר מקורי באנגליתarXiv:2609.11851v1 Announce Type: new Abstract: Language identification in code-mixed text, largely observed in social media, is highly essential when users frequently switch between multiple languages within a single utterance. Accurately identifying the languages of code-mixed tokens becomes an urgent necessity. Traditional language identification models, designed for monolingual text, are not well suited for token-level language identification in code-mixed settings. We formulate the task as a sequence labeling problem and fine-tune contextual transformer-based models MuRIL and XLM-RoBERTa best suited for Indian languages. We evaluate these systems on three different data configurations (Hindi, Gujarati, and Bengali) to predict language labels for individual tokens. We release a benchma
קרא במקור המקורי