כתבה
arXiv cs.CL ·
מודלים גדולים של שפה לסיווג תפקוד ציטוט
Large Language Models for Citation Function Classification
חוקרים בדקו מודלים שונים, כולל LLaMA, לסיווג תפקוד ציטוט. התוצאות הראו שיפור משמעותי ברמת הדיוק. המחקר גם הציג מאגר נתונים חדש, AC3, עם שבע קטגוריות לסיווג ציטוטים.
תקציר מקורי באנגליתarXiv:2607.17738v1 Announce Type: new Abstract: Citation function classification plays a crucial role in understanding the relationships between scientific publications and advancing bibliometric analysis. This study presents one of the first comprehensive evaluations of multiple state-of-the-art (SOTA) large language models (LLMs) for citation function classification, achieving new SOTA results on the ACL-ARC dataset. We systematically compare five models (Mistral 7B, Orca 2-7B, LLaMA 3.1-8B, Falcon 7B, and SciBERT) across zero-shot, few-shot, and fine-tuning approaches. Our fine-tuned Falcon 7B model achieves a 73.3% macro F1 score on ACL-ARC, representing a significant improvement over previous methods. Additionally, we introduce AC3, a novel dataset featuring a seven-category annotatio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית