יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

LLM Screening Performance: האם הביצועים נעצרו?

Has LLM Screening Performance Stalled in Software Engineering Systematic Reviews?
בדיקת ביצועי LLM במערכות סריקה מערכתית בהנדסת תוכנה: האם הביצועים נעצרו? המחקר חקר את ביצועי שמונה LLM חדשים, ומצא שהביצועים השתפרו קל, אך עדיין לא ניתן להחליף אותם בבני אדם.
תקציר מקורי באנגליתarXiv:2610.10633v1 Announce Type: cross Abstract: Screening in systematic reviews (SRs) is manual and time-consuming. Prior work has explored large language models (LLMs) for automating this step, but LLMs are evolving rapidly, so earlier performance claims may no longer accurately reflect their screening performance. We used an existing software engineering SR screening benchmark (SESR-Eval) as our data. We also power-sampled a new, smaller dataset (SESR-Eval-Mini) that allows evaluation at lower costs. Using this data, we evaluated eight new LLMs for screening performance. Additionally, we tested different prompts, analyzed LLM agreement in screening decisions and criteria, and examined the effect of refining the inclusion and exclusion criteria on screening performance. The eight new LL
קרא במקור המקורי