כתבה
arXiv cs.CL ·
בדיקת אחזור זיכרון בסוכנים קונברסאטיביים
When Users Don't Ask: Benchmarking Context-Driven Memory Retrieval in Conversational Agents
חוקרים פיתחו בדיקת אחזור זיכרון לסוכנים קונברסאטיביים. הבדיקה, LOCOMO-CONV, בודקת ארבע סגנונות שאילתות: דיאלוג, מרומז, נגדי ומורכב. הניסויים הראו פערים משמעותיים באחזור זיכרון בין בדיקות QA לבין שימוש קונברסאטיבי.
תקציר מקורי באנגליתarXiv:2609.03467v1 Announce Type: new Abstract: Large language models (LLMs) are increas- ingly deployed as long-horizon conversational agents, motivating growing interest in mem- ory systems. However, existing benchmarks primarily evaluate memory through QA-style probing rather than in-situ conversational usage. We introduce LOCOMO-CONV, a conversa- tional memory benchmark derived from Lo- CoMo with four query styles: dialog, implicit, counterfactual, and composed. Across five rep- resentative memory systems, we evaluate both retrieval recall and end-to-end response qual- ity. Our experiments show that conversational framing exposes substantial retrieval gaps over- looked by QA benchmarks, especially on im- plicit and composed queries, which multi-facet query rewriting narrows for raw-tur
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית