כתבה
arXiv cs.CL ·
Automatic register identification for the open web using multilingual deep learning
תקציר מקורי באנגליתarXiv:2406.19892v5 Announce Type: replace Abstract: This article presents multilingual deep learning models for identifying web registers -- text varieties such as news reports and discussion forums -- across 16 languages. We introduce the Multilingual CORE corpora, which contain over 72,000 documents annotated with a hierarchical taxonomy of 25 registers designed to cover the entire open web. Using multi-label classification, our best model achieves 79% F1 averaged across languages, matching or exceeding previous studies that used simpler classification schemes. This demonstrates that models can perform well even with a complex register scheme at multilingual scale. However, we observe a consistent performance ceiling across all models and configurations. When we remove documents with unc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית