כתבה
arXiv cs.LG ·
PowerBench: מדידת הטיה של מודלי שפה בבקשות של שינוי כוח
PowerBench: Measuring Language Model Bias in Power-shifting Requests
PowerBench: מערכת למדידת הטיה של מודלי שפה בבקשות של שינוי כוח. המחקר חוקר את ההבדלים בין המודלים בעזרת PowerBench, ומצא שהמודלים נוטים לסרב לבקשות של שינוי כוח כנגד פרטים.
תקציר מקורי באנגליתarXiv:2610.02303v1 Announce Type: new Abstract: Language models increasingly assist people with power-related requests, so systematic differences in whom they help could shift the distribution of power at scale, or be exploited by users who learn which identities are refused less. We introduce PowerBench, an evaluation of power-shifting requests that distinguishes self-empowerment, disempowerment, and power grabbing, plus a control of refusal-inducing requests that shift no power. We build, curate, and open-source a dataset of such requests varying the power domain, the context, the scale of the affected party, and the prior power standing of the user, and evaluate 24 models (12 from US and 12 from Chinese developers) under three experimental conditions: reciprocal nationalities of user an
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית