כתבה
arXiv cs.CL ·
E-CONAN: בנצ'מרקים לניבוי טקסטואלי ערבי
E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets
E-CONAN הוא בנצ'מרק חדש לניבוי טקסטואלי ערבי. הוא כולל זוגות משפטים ממקורות שונים, כולל תרגומים אוטומטיים ואימות אנושי. E-CONAN מורכב משני סטי נתונים: E-CONAN-2 ו-E-CONAN-3.
תקציר מקורי באנגליתarXiv:2609.11334v1 Announce Type: new Abstract: Natural Language Inference processes pairs of sentences to extract their semantic relations. NLI has been a hot research topic, integrated as a main component in other NLP applications. Despite significant advancements in textual inference across various languages all around the world, Arabic language still suffers from limited resources in this domain. To address this gap, this paper introduces E-CONAN benchmarks that are composed of sentences pairs from various sources: (1) automatically-translated pairs, (2) human-validated machine-translated pairs, (3) hand-crafted pairs from teaching Arabic as foreign language books, and (4) headlines pairs from different news channels containing rumors. E-CONAN contains two benchmark datasets, E-CONAN-2
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית