כתבה
arXiv cs.AI ·
A primer on evaluation methods for large language models in healthcare
תקציר מקורי באנגליתarXiv:2609.14819v1 Announce Type: cross Abstract: Large language models (LLMs) have a growing range of applications in medicine, and their evaluation is critical for ensuring they provide benefit and not harm. This evaluation can be more challenging than traditional machine learning for many reasons, including probabilistic and open-ended outputs, and behavior that shifts with prompt design and accumulated context. This review covers four key areas of LLM evaluation: principles of study design, statistical methods, capability evaluation and clinical context evaluation. Capability evaluation considers different benchmarks, including multiple-choice, agentic and multi-turn benchmarks, alongside operational metrics like token usage. Clinical context evaluation addresses establishing accuracy
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית