כתבה
arXiv cs.CL ·
VeriLLMed: Interactive Visual Debugging of Medical Large Language Models with Knowledge Graphs
תקציר מקורי באנגליתarXiv:2604.23356v2 Announce Type: replace Abstract: Large language models (LLMs) show promise in medical diagnosis, but real-world deployment remains challenging due to high-stakes clinical decisions and imperfect reasoning reliability. As a result, careful inspection of model behavior is essential for assessing whether diagnostic reasoning is reliable and clinically grounded. However, debugging medical LLMs remains difficult. First, developers often lack sufficient medical domain expertise to interpret model errors in clinically meaningful terms. Second, models can fail across a large and diverse set of instances involving different input types, tasks, and reasoning steps, making it challenging for developers to prioritize which errors deserve focused inspection. Third, developers struggl
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית