כתבה
arXiv cs.CL ·
Tracing mechanisms of sycophantic agreement in language models
תקציר מקורי באנגליתarXiv:2609.35822v1 Announce Type: new Abstract: Sycophantic agreement in language models refers to the tendency to overly affirm a user's stated beliefs or preferences, often at the expense of factual accuracy. Although it is widely recognized as an alignment failure, its underlying mechanisms remain poorly understood. In this work, we use causal mediation analysis to identify the mechanisms behind sycophantic agreement. We show that a stated opinion is incorporated into the residual stream of the final prompt token early, where it biases subsequent answer retrieval. A sparse set of early attention heads carries this opinion signal. Ablating these heads substantially reduces sycophancy while leaving factual accuracy largely intact. The same heads carry the opinion when it is explicitly sta
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית