כתבה
arXiv cs.CL ·
רובריק להערכת תהליכי חשיבה קליניים בתגובות מודלי שפה
A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
הוצע רובריק להערכת תהליכי חשיבה קליניים בתגובות מודלי שפה. הרובריק מבוסס על מסגרות הערכה בחינוך רפואי ומחקר קודם. הוא כולל קטגוריות להערכת תהליכי חשיבה ודיוק.
תקציר מקורי באנגליתarXiv:2609.37788v3 Announce Type: replace Abstract: We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR. BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability rules a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית