כתבה
arXiv cs.LG ·
IndicTalk: A Large-Scale Persona-Based Multilingual Conversational Corpus for Indic Languages
תקציר מקורי באנגליתarXiv:2607.23242v1 Announce Type: cross Abstract: Large Language Models (LLMs) have transformed conversational AI, yet high-quality multilingual code-mixed dialogue resources remain scarce, particularly for Indic languages where speakers naturally alternate between English and their native language in both native-script and Romanized forms. We present IndicTalk, one of the largest multilingual Indic code-mixed conversational corpora, comprising over 13,28,604 event-grounded multi-turn conversations across 18 language varieties covering 9 Indic languages. The corpus is generated through a fully automated pipeline that combines real-world news grounding, persona-conditioned dialogue generation using multilingual LLMs, and automatic quality validation. Extensive linguistic, automatic, and hum
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית