יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

הטיות עמוקות ורדודות במודלי שפה

Deep and shallow biases in language models
חוקרים מציגים שיטה למדידת הטיות במודלי שפה. הם מבדילים בין הטיות עמוקות להטיות רדודות, שתלויות בניסוח הפרומפט. המחקר מראה כי הטיות עמוקות קשות יותר להסרה מאשר הטיות רדודות.
תקציר מקורי באנגליתarXiv:2609.09901v1 Announce Type: new Abstract: Large language models often repeatedly select the same answer even when many alternatives are plausible. Prior work treats this concentration as bias, but it does not distinguish stable model preferences from responses that depend on a particular prompt wording. We introduce a bias depth score that measures both how strongly a model prefers its top answer under direct prompting and whether that answer survives scenario reframing. Across 4,442 opinion prompts and four large language models, only about a quarter of the concentrated preferences survive reframing. We call these persistent cases Deep biases, and the remaining prompt-dependent cases Shallow biases. Our results show that Deep biases are more often inherited from pretraining and pres
קרא במקור המקורי