כתבה
arXiv cs.AI ·
MEDIC: הערכה מקיפה של מדדים מובילים לבטיחות ויעילות LLM ביישומים קליניים
MEDIC: Comprehensive Evaluation of Leading Indicators for LLM Safety and Utility in Clinical Applications
MEDIC היא פלטפורמת הערכה מקיפה ל-Large Language Models (LLM) ביישומים קליניים. היא בודקת יכולות מעבר לשאילתות סטטיות, כולל ביצועים דטרמיניסטיים ו- Cross-Examination Framework (CEF). התוצאות מראות פערים משמעותיים בין יכולות סטטיות ליכולות פונקציונליות.
תקציר מקורי באנגליתarXiv:2409.07314v4 Announce Type: replace-cross Abstract: While Large Language Models (LLMs) achieve superhuman performance on standardized medical licensing exams, these static benchmarks have become saturated and increasingly disconnected from the functional requirements of clinical workflows. To bridge the gap between theoretical capability and verified utility, we introduce MEDIC, a comprehensive evaluation framework establishing leading indicators of clinical LLM competence across five dimensions. These upfront indicators reveal cross-benchmark capability gaps, such as the divergence between static knowledge retrieval and functional execution, that inform model selection before costly deployment-based evaluation. Beyond standard question-answering, we assess operational capabilities u
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית