יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

KlinikeBench: בדיקת מודלי שפה מעבר לדיוק אבחון

KlinikeBench: Evaluating Language Models Beyond Diagnostic Accuracy
KlinikeBench הוא בנץ'מרק לבדיקת מודלי שפה בסביבה קלינית. הוא כולל 333 משימות שנכתבו על ידי רופאים. מודלים כמו GPT ו-Claude נבדקו והראו פער גדול בין דיוק אבחון לביצועים במפגש קליני אינטראקטיבי.
תקציר מקורי באנגליתarXiv:2609.38480v1 Announce Type: cross Abstract: Most clinical benchmarks evaluate language models (LMs) on diagnosis using complete case descriptions. In clinical practice, however, patients present information in different ways, and clinicians must obtain relevant history and determine which examinations are needed before reaching a diagnosis. Diagnostic accuracy alone therefore cannot establish whether an agent gathered essential information or conducted an appropriate clinical assessment. Furthermore, existing benchmarks lack professional clinicians' verification. To address this gap, we introduce KlinikeBench, a benchmark of 333 clinician-authored tasks, each providing an isolated sandbox environment with a virtual patient, clinical tools, and task-specific success criteria. More tha
קרא במקור המקורי