יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

לא-וודאות אינה רשת סיוע ל-VQA קליני, אלא יכולה לצפות לכשל המודל?

Uncertainty Is Not a Safety Net for Clinical VQA, but Can It Anticipate Model Failure?
לא-וודאות במודלי VQA קליניים אינה מבטיחה בטיחות, אך יכולה לצפות לכשל המודל. המחקר חושף כי חלק מהמודלים נכשלים כאשר התשובה הנכונה נעלמת, ולא-וודאותם נשארת זהה. עם זאת, המחקר גם מצא כי לא-וודאות יכולה לצפות לכשל המודל.
תקציר מקורי באנגליתarXiv:2606.16583v2 Announce Type: replace Abstract: Safe deployment of clinical vision-language models (VLMs) requires reliable uncertainty estimation (UE): a signal indicating when predictions should be trusted or escalated to a clinician. We test whether current UE methods actually deliver this signal. Benchmarking 8 methods across 12 VLMs on clinical visual question-answering (VQA), we find that UE quality is not an intrinsic property of the UE method: it tracks model accuracy, degrading precisely where the model performance is weakest, and therefore where reliability is most needed. When we stress-test models by hiding the correct option among the multiple-choice answers (NOTA perturbations), accuracy collapses while uncertainty barely changes, leaving models systematically miscalibrat
קרא במקור המקורי