כתבה
arXiv cs.AI ·
כאשר זיכרון עוזר? בדיקה עם חשבון עלויות של זיכרון ארוך-טווח ב-LLM Agents
When Does Memory Help? A Cost-Aware Evaluation of Long-Term Memory in Tool-Using LLM Agents
בדיקה של זיכרון ארוך-טווח ב-LLM Agents עם חשבון עלויות
תקציר מקורי באנגליתarXiv:2609.05441v1 Announce Type: new Abstract: Long-term memory for LLM agents is evaluated today by conversational recall benchmarks (LoCoMo, LongMemEval), which measure question answering over dialogue history, not whether remembered facts change what a tool-using agent does. We present MERIT (Memory Evaluation for Realistic Instrumented Tasks), a benchmark and harness that measures the marginal utility of memory for task-executing agents under explicit cost accounting. MERIT provides episodic tool-use tasks in three domains whose dependence on earlier-episode facts is verified by an automated leak check; a difficulty ladder ending in updated-fact recall; controlled memory corruption; and full token and dollar metering of every memory operation. Across 23,440 scored episodes ($42.57), a
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית