יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

זיהוי תפקיד של פסקאות בשפה הבנגלית: פיתוח קורפוס, בדיקה של מודלים והבהרה

Bangla Sentence Function Classification: Corpus Development, Model Benchmarking, and Interpretability
פיתוח קורפוס ובדיקה של מודלים לזיהוי תפקיד של פסקאות בשפה הבנגלית. הקורפוס כולל 10,000 פסקאות, מאותויים ידנית לארבע קטגוריות תפקודיות. המחקר כולל גם בדיקה של מודלים שונים, כולל רשתות עצבים ומודלי רכיבי רשת. המחקר גם כולל הבהרה של המודלים, כדי להבהיר את ההחלטות שלהם.
תקציר מקורי באנגליתarXiv:2609.13869v1 Announce Type: cross Abstract: Automatic sentence function identification is important for many downstream natural language processing (NLP) applications such as dialogue systems, text-to-speech synthesis, and machine translation. However, benchmark resources for Bangla sentence function classification remain limited. To mitigate this gap, this paper introduces a corpus of 10,000 Bangla sentences, manually annotated into four functional categories, namely declarative, interrogative, imperative, and exclamatory. The corpus is nearly balanced across the four classes, with high annotation reliability reflected by a Fleiss\' Kappa of 0.82. Furthermore, we evaluate multiple feature representations, including Bag-of-Words (BoW), TF-IDF, and Word2Vec, with several classical mac
קרא במקור המקורי