כתבה
arXiv cs.AI ·
הגנה-בעומק ל-LLMs: בדיקת שערי זיכרון נגד סימפתיה-עוררת-זיכרון
Defense-in-Depth for LLMs: Evaluating Memory Gates Against Activation-Induced and Memory-Induced Sycophancy
בדיקת הגנות נגד סימפתיה-עוררת-זיכרון ב-LLMs. המחקר חוקר את השפעת שערי זיכרון על תוצאות LLMs ומציע פתרונות להגברת אמינותם.
תקציר מקורי באנגליתarXiv:2610.07403v1 Announce Type: new Abstract: Long-term memory allows Large Language Models (LLMs) to maintain personalized context across interactions, but retrieved user history can induce memory-induced sycophancy, causing models to favor stored user beliefs over objective evidence. Existing defenses primarily operate on retrieved context and are rarely evaluated jointly with internal behavioral bias. We introduce a $2 \times 2$ defense-in-depth framework separating internal activation steering from external memory handling. We extract sycophancy steering directions from 100 paired prompts and evaluate four open-weight models across 10 steering coefficients and five memory-defense configurations on MemSyco-Bench (answers for all 1,550 items; defense conditions judged on a fixed 250-it
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית