כתבה
arXiv cs.CL ·
FairFund-Bench: בדיקת גישה חודשה למודלי LLM לקביעת חומרה
FairFund-Bench: Evaluating Distributive Bias in LLM Resource Allocation
FairFund-Bench: בדיקת גישה חודשה למודלי LLM לקביעת חומרה. הבדיקה נועדה לבדוק כיצד מודלי LLM קביעים כיצד להפצת משאבים ספורדיים, ואם הם נוטים לפגוע בקבוצות חלשות. הבדיקה כוללת 600 בקשות לסיוע כספי, שנוצרו באמצעות דגמי אנושיים, ובוחן 14 מודלי LLM שונים.
תקציר מקורי באנגליתarXiv:2607.28934v2 Announce Type: replace Abstract: Large language models (LLMs) are increasingly involved in the distribution of scarce resources, raising concerns about biased allocations based on characteristics like race and gender. Recent LLM audits have produced inconsistent results, however, finding evidence of both positive and negative discrimination towards women and ethnic minorities, even for the same models. We show that this disagreement can arise from differences in audit format and introduce FairFund-Bench, a benchmark that systematically varies key features of previous audit designs: the evaluation task (rating, ranking, or allocation), comparison context (single or multi-stimulus), and whether the audit is transparent or disguised. The benchmark comprises 600 English-lang
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית