כתבה
arXiv cs.AI ·
A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
תקציר מקורי באנגליתarXiv:2609.37788v2 Announce Type: replace-cross Abstract: We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR.BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includes provisional behavioural anchors, applicability ru
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית