כתבה
arXiv cs.LG ·
התבדלות הייצוגים הפנימיים של סיקופנטיות ב-LLMs
Dissociating the Internal Representations of Sycophancy in LLMs
חוקרים התבדלו את הייצוגים הפנימיים של סיקופנטיות ב-LLMs ומצאו שהם מיוצגים באופן שונה ב-LLMs שונים.
תקציר מקורי באנגליתarXiv:2607.07003v2 Announce Type: replace Abstract: Large Language Models (LLMs) frequently exhibit sycophancy, agreeing with a user's statement even when it is incorrect. While often studied as a single, uniform behavior, sycophancy can manifest in substantially distinct ways across contexts, raising the question of whether this heterogeneity is reflected in its internal mechanisms. To address this gap, we dissociate the representations of sycophancy into factual and opinion subtypes, motivated by prior evidence of heterogeneous truth representations in LLMs. We train linear probes and construct steering vectors on one subtype's activations and evaluate their transfer to the other, measuring the extent to which representations are shared and visualizing them via Linear Discriminant Analys
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית