יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הפער הגדל בבדיקות השכיחות של מודלי שפה גדולים במחקר רפואי 2023-2026

The widening evaluation gap in medical large language model research 2023 to 2026
במחקרי מודלי שפה גדולים רפואיים, הפער בין הבדיקות השכיחות לבין המודלים החדשים גדל. 62% מהניסויים המסודרים באופן רנדומלי בחנו מודלים שהיו כבר נגמרים.
תקציר מקורי באנגליתarXiv:2609.11770v1 Announce Type: new Abstract: Large language models are superseded every few quarters; clinical evidence takes years. We asked whether medical research is keeping pace with the systems it evaluates. PubMed returned 11,628 records for January 2023 to June 2026 across fourteen clinical domains, growing 45-fold; 2.5% used a randomised, controlled or prospective design. Evaluation lag, from a study's newest named model release to its own publication, widened from 1.33 to 6.08 quarters. Because discontinued models age mechanically, we benchmarked this against a counterfactual holding model composition fixed: migration to newer systems offset only 56% of the drift (95% CI 50-65). Randomised trials evaluated models a median 4.6 quarters older than other designs (P = 3 x 10^-19),
קרא במקור המקורי