כתבה
arXiv cs.AI ·
מקור הכשלים בניתוח LLM
Where Do Apparent LLM Clinical Triage Failures Arise? Localizing the Multiple-Choice Format Effect
חוקרים בדקו מדוע מודלים LLM נכשלים בניתוח רפואי. הם מצאו שהבעיה נובעת מהפורמט הרב-ברירתי. המחקר השתמש במודל Qwen.
תקציר מקורי באנגליתarXiv:2605.29889v2 Announce Type: replace-cross Abstract: LLM evaluations using clinician-authored triage vignettes have reported substantial under-triage under constrained multiple-choice testing. Yet model performance on the same clinical cases can change when responses are generated in free text. We test whether this format effect appears while the case is processed or when clinical information is mapped to the final answer. Using sparse-autoencoder (SAE) features in Gemma 3 4B/12B IT and Qwen3-8B, we find that medical features fire on the shared clinical narrative under both formats but are inactive at the multiple-choice decision token. Emergency-tier information is linearly decodable from vignette representations with ROC-AUC $0.95$--$1.00$ under both formats, with no significant for
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית