כתבה
arXiv cs.CL ·
BTBR: פלטפורמה בייסית-תאורטית להסרת הטיה אינטואיטיבית ממודלי שפה גדולים
BTBR: A Bayesian-Theory-Driven Probabilistic-Fuzzy Framework for Implicit Bias Removal in Large Language Models
פלטפורמה להסרת הטיה אינטואיטיבית ממודלי שפה גדולים. BTBR משתמשת בבינה מלאכותית ובאלגוריתמים פלסטיים כדי לזהות ולהסיר הטיה אינטואיטיבית ממודלי שפה. הפלטפורמה נבחנה במספר מקרים והראתה תוצאות טובות.
תקציר מקורי באנגליתarXiv:2408.10608v2 Announce Type: replace Abstract: Large language models (LLMs) may encode biased associations from heterogeneous training corpora that are not immediately visible under ordinary prompting, but can surface when the model is steered toward particular demographic personas. Such behavior often manifests not as explicit toxic output, but as systematic performance differences across semantically equivalent tasks, making the resulting bias difficult to detect and mitigate. To address this issue, we formalize the implicit bias problem as persona-induced performance disparity and argue that bias evidence should be treated as a graded signal rather than a binary label. Motivated by this observation, we model biased knowledge as a fuzzy subset equipped with an explicit membership fu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית