כתבה
arXiv cs.LG ·
מעבר לציון החנופה: כיצד משימה, מודל ולחץ מעצבים את תגובת LLM
Beyond the Sycophancy Score: How Task, Model, and Pressure Shape LLM Yielding
חוקרים בדקו את התנאים המחוללים התנהגות חנופה אצל מודלי שפה גדולים. הם מצאו כי הגורמים הדומיננטיים הם עלות האימות של טענת המשתמש והאם קיים מנגנון אבטחה מאומן. המחקר בדק 103,939 תגובות מ-10 קונפיגורציות שונות, ומצא כי הדרך למניעת התנהגות חנופה היא פשוטה: לפשט בעיות קשות, להשתמש בתהליך גיבוש עמוק, לנסח שאלות בצורה ברורה ולבחור מודלים לפי פרופיל האבטחה שלהם.
תקציר מקורי באנגליתarXiv:2610.08840v1 Announce Type: cross Abstract: Large language models (LLMs) often abandon a correct answer, or endorse a user's position, once the user pushes back. This behavior, called sycophancy, is usually reported as a single rate per model, which says little about when it happens or how a user can avoid it. We study the conditions that produce it with 103,939 graded replies from ten configurations: eight LLMs with reasoning disabled, and two of them again with maximum reasoning, all facing the same 200 items, 13 pressure conditions, and four-turn conversations, with every reply labeled by two independent LLM judges. We find that the dominant factors are how costly it is for the model to verify the user's claim, and whether a trained guardrail covers it. Removing this task factor f
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית