יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מעבר לצללים של מערת פלאטו

Beyond the Shadows of Plato's Cave: Evaluating False Memory in Autonomous Agents via Counterfactual Reasoning
חוקרים פיתחו שיטה להערכת 'זיכרון כוזב' בסוכנים אוטונומיים. השיטה, FAME, משתמשת בתרחישים נגדיים כדי לבדוק כיצד האמונות של הסוכנים משתנות תחת השפעות היפותטיות. הניסויים הראו כי FAME מצליחה לזהות 'זיכרון כוזב' ברמת דיוק גבוהה.
תקציר מקורי באנגליתarXiv:2609.39473v1 Announce Type: new Abstract: Autonomous agents increasingly rely on memory to generalize beyond their training environments. However, agents are bounded by what they have seen and believed, and leveraging such memories in unseen environments can introduce biases into their internal beliefs. We formalize this phenomenon as \textit{false memory}, which can arise from spurious correlations, environment shifts, and knowledge conflicts. Despite its importance, false memory is difficult to evaluate because it stems from agent internal beliefs and is easily confounded with ordinary generalization failures. Therefore, we propose FAME, a training-free framework that evaluates false memory through the evolution of agent beliefs under counterfactual reasoning. Specifically, counter
קרא במקור המקורי