כתבה
arXiv cs.CL ·
Alignment Reduces Expressed but Not Encoded Gender Bias: A Unified Framework and Study
תקציר מקורי באנגליתarXiv:2603.24125v3 Announce Type: replace Abstract: During training, Large Language Models (LLMs) learn social regularities that can lead to gender bias in downstream applications. Most mitigation efforts focus on reducing bias in generated outputs, typically evaluated on structured benchmarks, which raises two concerns: output-level evaluation does not reveal whether alignment modifies the model's underlying representations, and structured benchmarks may not reflect realistic usage scenarios. We propose a unified framework to jointly analyze intrinsic and extrinsic gender bias in LLMs using identical neutral prompts, enabling direct comparison between gender-related information encoded in internal representations and bias expressed in generated outputs. Contrary to prior work reporting we
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית