כתבה
arXiv cs.AI ·
הצעת רושם לביקורת סביב תגובות של מודלי שפה גדולים לקליניקה
A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
הצעת רושם לביקורת סביב תגובות של מודלי שפה גדולים לקליניקה. הרושם נועד לבחון את התגובות של המודלים לקליניקה ולבדוק את היכולת שלהם להציג תגובות קליניות.
תקציר מקורי באנגליתarXiv:2609.37788v3 Announce Type: replace-cross Abstract: We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR. BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability r
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית