כתבה
arXiv cs.LG ·
MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents
תקציר מקורי באנגליתarXiv:2609.24259v3 Announce Type: replace Abstract: The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group relative
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית