יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

שאלות ותשובות למסמכי נפסרט: קבוצת נתונים נמוכת-משאבים ליישומי שירות ציבורי

Nepali Passport Question Answering: A Low-Resource Dataset for Public Service Applications
במחקר זה, נפתחה קבוצת נתונים חדשה לשאלות ותשובות בנפאלי, עבור יישומי שירות ציבורי. המחקר כולל טיפול במודלי SBERT ו-E5, ובחינת תוצאותיהם.
תקציר מקורי באנגליתarXiv:2603.13320v2 Announce Type: replace-cross Abstract: Nepali, a low-resource language, faces significant challenges in building an effective information retrieval system due to the unavailability of annotated data and computational linguistic resources. In this study, we attempt to address this gap by preparing a pair-structured Nepali Question-Answer dataset. We focus on Frequently Asked Questions (FAQs) for passport-related services, building a data set for training and evaluation of IR models. In our study, we have fine-tuned transformer-based embedding models for semantic similarity in question-answer retrieval. The fine-tuned models were compared with the baseline BM25. In addition, we implement a hybrid retrieval approach, integrating fine-tuned models with BM25, and evaluate the
קרא במקור המקורי