יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

קרשי יציבות: הבחנה בין כיווץ לרעש בבדיקות צדקנות של סוכני LLM קליניים עם FairMedAgent

Instability Floors: Separating Bias from Noise in Fairness Audits of Clinical LLM Agents with FairMedAgent
ניטרול קרשי יציבות בסוכני LLM קליניים: פיתוח FairMedAgent
תקציר מקורי באנגליתarXiv:2609.03221v3 Announce Type: cross Abstract: Counterfactual fairness audits of clinical language-model agents report a flip rate: how often an action changes when only the patient's demographic descriptor changes. Part of that rate is not demographic. A stochastic agent also changes its own action when nothing changes, and a flip rate cannot be interpreted without knowing how often. We measured it. Re-running one condition ten times over sixteen synthetic vignettes at default sampling changed a clinical agent's action in 8.7 percent of replicate pairs, from 2.2 percent for intensive-care escalation to 17.9 percent for controlled-substance caution, an output given no operational criteria. Across six models from five vendors, pooled floors ranged from 2.5 to 23.7 percent; in this panel
קרא במקור המקורי