יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בדיקת מודלים גדולים לשאלות רפואיות פתוחות

Evaluating Large Language Model Raters for German Open-Response Clinical Questions: A Physician-Annotated Benchmark Study of Agreement, Evaluator Bias, and Abstention
חוקרים בדקו את MedQADE, בנך' גרמני לשאלות רפואיות פתוחות. הם השוו את התוצאות של מודלים שונים, כולל Gemini, עם תשובות של רופאים. התוצאות הראו שמודלים חזקים יכולים להתקרב לרמת ההסכמה של רופאים, אך עדיין יש צורך באימות רפואי.
תקציר מקורי באנגליתarXiv:2607.01103v3 Announce Type: replace Abstract: Background: Expert-annotated benchmarks for non-English open-response clinical questions are scarce. LLM-as-a-judge systems may scale evaluation but require validation. Objective: To introduce MedQADE, a standardized German open-response clinical benchmark with physician reference annotations, and evaluate LLM-as-a-judge alignment, self- and intra-family bias, and abstention. Methods: The benchmark contains 3,800 question-answer sets with answers from five student LLMs and annotations from 10 physicians. All 10 rated the 200-question core; two primary raters assessed each of 3,600 extension questions, with the tenth resolving disagreements. Nine LLM evaluators assessed all sets. We assessed physician reliability, student-model accuracy, e
קרא במקור המקורי