כתבה
arXiv cs.CL ·
אימון ראשון: כלי לבחינת זיכרון של סוכנים
Ground Truth First: A Longitudinal Evaluation Instrument for Agent Memory, and the Tenure Crossover in Memory-Architecture Rankings
חוקרים פיתחו כלי לבחינת זיכרון של סוכנים. הכלי בוחן את היכולת של סוכנים לזכור מידע לאורך זמן. הניסויים הראו שהדירוגים של הסוכנים משתנים עם הזמן.
תקציר מקורי באנגליתarXiv:2607.21962v1 Announce Type: new Abstract: Benchmarks for LLM-agent memory typically generate conversations first and extract answer keys afterwards -- with documented label-error and contamination problems -- and they overwhelmingly measure short interaction histories. We invert the pipeline: a seeded life-script sampler emits facts with validity intervals, volatility classes, and source channels before any text exists; an LLM renderer writes chat and email from per-event fact manifests; a fidelity verifier confirms every planted fact; and questions are instantiated mechanically from the script, so gold answers are script-valid by construction and separately validated for answerability. The synthetic, fictionalized corpus (~380 questions, 15 types) embeds features absent from the ben
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית