יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

רובריק להערכת תהליכי חשיבה קליניים בתגובות מודלי שפה

A Proposed Rubric for Evaluating Expressed Clinical Reasoning in Large Language Model Responses
חוקרים הציעו רובריק להערכת תהליכי חשיבה קליניים בתגובות מודלי שפה. הרובריק מבוסס על מסגרות הערכה בחינוך רפואי ומחקר קודם. הוא כולל קטגוריות להערכת תהליכי חשיבה ודיוק.
תקציר מקורי באנגליתarXiv:2609.37788v1 Announce Type: new Abstract: Rubrics support the structured evaluation of language models. We propose a rubric for assessing expressed clinical reasoning in model responses, drawing on three bodies of work: medical education assessment frameworks (ART, SCT, Key Feature Problems and OSCE); clinical LLM benchmarks (MedR-Bench, HealthBench, TIMER-Bench, DR.BENCH, PrIME-LLM and PatientSafeBench); and general LLM reasoning evaluation research, including the Factuality-Validity-Coherence-Utility taxonomy, FaithCoT-Bench and C2-Faith. We use groundedness as a clinically oriented adaptation of the taxonomy's factuality category. The rubric brings these concepts together in a multidimensional framework for scoring free-text responses to gold-standard clinical vignettes. It includ
קרא במקור המקורי