יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

AcuityBench: בדיקת הערכת חומרה קלינית ואי-הוודאות

AcuityBench: Evaluating Clinical Acuity Identification and Uncertainty Alignment
AcuityBench - בדיקת הערכת חומרה קלינית ואי-הוודאות. המאמר עוסק בפיתוח AcuityBench, תקן לבדיקת הערכת חומרה קלינית ואי-הוודאות. התקן כולל 914 מקרים, כולל 697 מקרים שנקבעו באופן סטנדרטי לבדיקת דיוק ו-217 מקרים שנקבעו באופן רפואי לבדיקת אי-הוודאות. AcuityBench תומך בשני פורמטים של תפקידים: סיווג רב-מוטציה בסביבת QA ותגובות קונברציונליות חופשיות שנבדקות על ידי יוצר.
תקציר מקורי באנגליתarXiv:2605.11398v2 Announce Type: replace-cross Abstract: We introduce AcuityBench, a benchmark for evaluating whether language models identify the appropriate urgency of care from user medical presentations. Existing health benchmarks emphasize medical question answering, broad health interactions, or narrow workflow-specific triage tasks, but they do not offer a unified evaluation of acuity identification across these settings. AcuityBench addresses this gap by harmonizing five public datasets spanning user conversations, online forum posts, clinical vignettes, and patient portal messages under a shared four-level acuity framework ranging from home monitoring to immediate emergency care. The benchmark contains 914 cases, including 697 consensus cases for standard accuracy evaluation and
קרא במקור המקורי