כתבה
arXiv cs.CL ·
MyoCardBench: בנך' לבדיקת מודלי שפה גדולים בסיעורי לב
MyoCardBench: A Real-World Data Benchmark for Evaluating Large Language Models in Clinically Authentic Cardiovascular Care Scenarios
MyoCardBench הוא בנך' לבדיקת מודלי שפה גדולים בסיעורי לב. הבנך' כולל 2,263 פריטים מ-13 מאגרי נתונים. GPT-5.4 השיג את התוצאה הטובה ביותר.
תקציר מקורי באנגליתarXiv:2607.25186v1 Announce Type: new Abstract: Background: Most medical large language model (LLM) benchmarks focus on examination knowledge or isolated tasks and may not reflect the longitudinal, multimodal, and safety-critical workflow of cardiovascular care. Objective: To develop MyoCardBench, a real-world benchmark spanning the cardiovascular care continuum, and assess LLM performance across clinical dimensions and specialist tasks. Methods: MyoCardBench includes 2,263 items from 13 task-specific datasets derived from de-identified cardiovascular records and examination data. Sixteen cardiology physicians conducted annotation and reference construction, followed by cross-review from two senior cardiologists. Seven LLMs generated 15,841 outputs under standardized zero-shot settings. Op
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית