כתבה
arXiv cs.AI ·
שאילתא תגיות והיפוך דורות של סיכופנטיות ב-45 מודלים
Tag Questions and the Generational Reversal of Sycophancy Across 45 Language Models
חוקרים בדקו את השפעת שאילתא תגיות על 45 מודלים, כולל GPT, Claude ו-Qwen. התוצאות הראו היפוך דורות של סיכופנטיות, כאשר מודלים חדשים יותר נוטים להתנגד לבחירות. המחקר מראה כי התופעה קשורה למבנה השטחי של השאילתא ולא לעמדת המשתמש.
תקציר מקורי באנגליתarXiv:2607.23976v1 Announce Type: cross Abstract: Appending a two-word confirmation tag to a decision question -- "Is X the better choice?" versus "X is the better choice, right?" -- changes whether a language model endorses the choice. We measure this tag effect on 20 frozen, ground-truth-free decisions between two defensible options, counterbalanced so a model's own preferences cancel, scored by exact match on clamped yes/no replies -- no LLM judge, no embeddings. Across 45 models the effect spans +32% to -32% -- a 64-point swing on one word -- with 5 models significantly sycophantic and 17 significantly resistant (BH-FDR q=.10). The sign is a clock: within model families the effect crosses from positive to negative as generations advance (GPT +4 to -28; Claude +7 to -32; Qwen and Grok l
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית