יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

מסגרת ביצועים לאוטומציה של בדיקות מערכתיות

A Benchmark Framework for Screening Automation in Systematic Reviews
מחקר זה מציג מסגרת ביצועים להערכת ביצועי מודלים של LLM בבדיקות מערכתיות. המסגרת כוללת מאגר נתונים של 45,064 כניסות מתויגות וכלי תמיכה לניסויים וניתוח תוצאות.
תקציר מקורי באנגליתarXiv:2609.30298v1 Announce Type: new Abstract: Systematic reviews (SR) are essential for evidence-based research, but their screening phase is highly time-consuming and labor-intensive. Large language models (LLMs) offer a promising opportunity to reduce this workload by assisting with article relevance classification. However, existing evaluation approaches often rely on traditional metrics that may be misleading for highly imbalanced SR screening datasets.This paper presents a benchmark dataset of $45\,064$ labeled entries for evaluating LLM performance in SR screening across 32 curated secondary studies. It proposes an evaluation framework that accounts for class imbalance, i.e., the natural prevalence of excluded articles relative to included articles in SRs. It also introduces Prompt
קרא במקור המקורי