כתבה
arXiv cs.CL ·
הטיה של סמכות במודלים של שפה
Authority Bias in Language Models: Source Deference and User Agreement Are Not Interchangeable
חוקרים גילו הטיה של סמכות במודלים של שפה, כאשר מודלים מסכימים יותר עם טענות המיוחסות למקורות מהימנים. המחקר בדק חמש משפחות של מודלים ומצא שהוספת הערת 'מקור מהימן' לטענה שגויה גרמה ל-45-88% מהמודלים לשנות את תשובתם.
תקציר מקורי באנגליתarXiv:2609.37616v1 Announce Type: cross Abstract: Language models tend to agree with whatever a user asserts, and post-training increasingly targets this sycophancy so that models evaluate claims on their merits rather than deferring to the user. Yet the same models are far more compliant when a wrong answer is attributed to a verified source, which is how retrieval results, tool outputs, and grounded-search content often present information. We measure this gap across five open-weight families and three closed APIs. A single verified-source note endorsing a wrong answer flips 45-88% of baseline-correct responses in seven of eight models, and compliance rises with how authoritative the note sounds. Source deference and user agreement are not behaviorally interchangeable inside the model: o
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית