כתבה
arXiv cs.CL ·
PReSS: An Automated Black-Box Framework for Evaluating Political Stance Stability in LLMs
תקציר מקורי באנגליתarXiv:2504.17052v4 Announce Type: replace Abstract: Existing evaluations of political bias in large language models (LLMs) typically classify outputs as left- or right-leaning. We extend this perspective by examining how ideological tendencies vary across topics and how consistently models maintain their positions, a property we refer to as stability. To capture this dimension, we propose PReSS (Political Response Stability under Stress), an automated black-box framework that evaluates LLMs by jointly considering model and topic context, categorizing responses into four stance types: stable-left, unstable-left, stable-right, and unstable-right. Applying PReSS to 9 widely used LLMs across 19 political topics reveals substantial variation in stance stability; for instance, a model that is le
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית