כתבה
arXiv cs.CL ·
CHI-Bench: אוטומציה של תהליכי בריאות
CHI-Bench: Can AI Agents Automate End-to-End, Long-Horizon, Policy-Rich Healthcare Workflows?
CHI-Bench הוא בנץ'מרק לאוטומציה של תהליכי בריאות. הוא בודק יכולתם של סוכנים לבצע משימות מורכבות בתחום הבריאות, כולל אישורים רפואיים וניהול שימוש. התוצאות מראות כי הסוכנים הטובים ביותר מצליחים לבצע רק 28% מהמשימות.
תקציר מקורי באנגליתarXiv:2605.16679v3 Announce Type: replace Abstract: End-to-end automation of realistic healthcare operations stresses three capabilities underrepresented in current benchmarks: policy density, decisions must be grounded in a large library of medical, insurance, and operational rules; Multi-role composition: a single task requires the agent to play multiple roles with handoffs; and multilateral interaction: intermediate workflow steps are multi-turn dialogs, such as peer-to-peer review and patient outreach. We introduce $\chi$-Bench, a benchmark of long-horizon healthcare workflows across three domains: provider prior authorization, payer utilization management, and care management. Each task hands the agent a clinical case in a high-fidelity simulator of 20 healthcare apps exposed via 87 M
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית