יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

ביקורת על יציבות ואמינות של LLMs באבחון רפואי

The Reliability of LLMs for Medical Diagnosis: An Examination of Consistency, Manipulation, and Contextual Awareness
במחקר זה נבחנה יציבות ואמינות של LLMs באבחון רפואי. נמצא כי שני המודלים, Gemini ו-ChatGPT, הציגו יציבות 100% באבחון, אך היו חששות לגבי רגישותם לפרטים לא רלוונטיים ולשינויים בהקשר.
תקציר מקורי באנגליתarXiv:2503.10647v2 Announce Type: replace-cross Abstract: This study evaluated the diagnostic reliability of two Large Language Models (LLMs), Google Gemini 2.0 Flash and OpenAI ChatGPT-4o, across three dimensions: consistency under rephrased inputs, susceptibility to irrelevant prompt content, and responsiveness to added clinical context. We designed 52 clinical scenarios and modified each under controlled conditions. For consistency, scenarios were rephrased with demographic, wording, and examination changes that preserved the diagnostic core. And the susceptibility was evaluated through embedding irrelevant but plausible narrative details while keeping the clinical evidence unchanged. For contextual awareness, patient history, lifestyle data, or diagnostic findings were added to shift t
קרא במקור המקורי