כתבה
arXiv cs.CL ·
הסטייה האפיסטמית של מדיניות ב-LLM
Epistemic Policy Divergence in Multi-Turn LLM Contamination: A Protocol-Gradient Investigation
חוקרים בדקו את ההשפעה של הכנסת ידיעות שגויות לשיחות LLM. הם מצאו ש-GPT-5.4 Mini לא אימץ ידיעות שגויות, בעוד Gemini-3.1 Flash-Lite הראה נטייה חדה לאימוץ ידיעות שגויות כאשר הן הוצגו כמקורות מהימנים.
תקציר מקורי באנגליתarXiv:2609.35308v2 Announce Type: replace Abstract: Large language models treat conversation history as unverified context, so false premises injected into prior turns can be adopted as fact, a failure mode we term session-level contamination. We introduce five contamination protocols arranged along a source-authority gradient, holding the false premise constant while varying its epistemic framing, and evaluate GPT-5.4 Mini, Gemini-3.1 Flash-Lite, and GLM-4.5-Air across ten knowledge domains at temperature zero (22,500 turns), judged by a dual-track automated evaluator validated against a human gold standard (Cohen's kappa = 1.000 for binary adoption; 0.92 linear-weighted for collapse severity). GPT-5.4 Mini recorded zero adoptions across all 500 sessions; a base-model logit probe shows it
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית