כתבה
arXiv cs.CL ·
אותו מטופל, הזמנה שונה
Same Patient, Different Order: Action-Level Reliability of Clinical LLM Agents Under Repeated Runs
סוכנים רפואיים מבוססי LLM מראים הבדלים ברמת הפעולה. מחקר חדש בודק את האמינות של סוכנים אלו בהזמנות חוזרות. התוצאות מראות שהסוכנים יכולים להפיק תוצאות שונות עבור אותן קלטים.
תקציר מקורי באנגליתarXiv:2609.13582v1 Announce Type: new Abstract: A clinical agent benchmark can report the same verdict on identical inputs while the agent files a materially different order on each run. Such agents order tests, request medications and place referrals, yet benchmarks typically score one run per task and rarely ask whether identical inputs produce identical actions; MedAgentBench, the benchmark we use, scores a single attempt and says so. To measure this gap we introduce "same-input rerun", which replays a task with every input held fixed and compares the orders rather than the score, with six reliability metrics, and apply it to 1000 MedAgentBench runs across 50 tasks from its five write-capable families, two open-weight models below ten billion parameters quantised to four bits, and two t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית