כתבה
arXiv cs.CL ·
AtlasNLP: אטלס מודע למדינות של ייצוגי מאגרי נתונים ב-NLP
AtlasNLP: A Country-Aware Atlas of Dataset Representation in NLP
AtlasNLP הוא אטלס המתעד יותר מ-13,000 רשומות של מאגרי נתונים ב-NLP, ובוחן את הייצוג הגאוגרפי שלהם. המחקר מראה כי ישנם פערים גדולים בייצוג הנתונים בין מדינות ומטלות שונות.
תקציר מקורי באנגליתarXiv:2608.30107v2 Announce Type: replace Abstract: Understanding which countries are represented in NLP datasets is essential for identifying gaps, targeting data collection, measuring progress, and informing AI policy. However, geographic metadata is very rarely available, and country-level representation is often hidden behind broad language-level claims. We introduce AtlasNLP, a country-aware atlas of over 13,000 NLP dataset records across normalized NLP task categories, tracking both the populations represented and where datasets are produced. AtlasNLP includes AtlasNLP-Gold, a human-curated reference set, and AtlasNLP-Core, an ACL-derived large-scale collection. Using this resource, we show that (1) dataset coverage is highly uneven across countries and tasks; (2) dataset production
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית