יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

סימפתיה אינה דבר אחד: הפרדה סיבתית של התנהגויות סימפתיות ב-LLMs

Sycophancy Is Not One Thing: Causal Separation of Sycophantic Behaviors in LLMs
במאמר זה, החוקרים מפרידים בין התנהגויות סימפתיות של LLMs לבין התנהגויות אמיתיות. הם מצאו שהתנהגויות אלה קשורות למודלים שונים ושיטות שונות. המחברים טוענים שהתגליות שלהם יכולות לשפר את הבינה העמוקה של LLMs.
תקציר מקורי באנגליתarXiv:2509.21305v4 Announce Type: replace Abstract: Large language models (LLMs) often exhibit sycophantic behaviors -- such as excessive agreement with or flattery of the user -- but it is unclear whether these behaviors arise from a single mechanism or multiple distinct processes. We decompose sycophancy into sycophantic agreement and sycophantic praise, contrasting both with genuine agreement. Using difference-in-means directions, activation additions, and subspace geometry across multiple models and datasets, we show that: (1) the three behaviors are encoded along distinct linear directions in latent space; (2) each behavior can be independently amplified or suppressed without affecting the others; and (3) their representational structure is consistent across model families and scales.
קרא במקור המקורי