כתבה
arXiv cs.AI ·
MemCalib: תקן לבדיקה ואופטימיזציה של שימוש במסגרת זיכרון ב-LLM
MemCalib: Benchmarking and Optimizing Memory Use in LLM Agents
במאמר זה, MemCalib, נציגים תקן לבדיקה ואופטימיזציה של שימוש במסגרת זיכרון ב-LLM. התקן נבדק על קבוצת ניסויים והתוצאות הראו שדגמי LLM קיימים נכשלו להשתמש בזיכרון באופן ראוי. המחברים מציעים תקן חדש, MemCalib-RL, שמסוגל להשתמש בזיכרון באופן ראוי ולשפר את התוצאות.
תקציר מקורי באנגליתarXiv:2609.24259v3 Announce Type: replace-cross Abstract: The effectiveness of agent memory ultimately depends on whether the underlying LLM gives each memory in context an appropriate degree of influence over its response. Yet this capability has remained largely overlooked. To assess this capability, we introduce MemCalib, a benchmark grounded in realistic memory-system scenarios for evaluating memory use and advancing optimization algorithms. Results on the MemCalib test set reveal that frontier open- and closed-source models struggle to use memory appropriately. They frequently over-use or under-use memory rather than matching each proposition's actual use to its target level, leading to biased, low-quality responses. Experiments with common post-training algorithms, including group re
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית