יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

מצב האמונה מאורגן והבכורה של Precision-Aware Benchmark ל-LMM Memory Retrieval

Structured Belief State and the First Precision-Aware Benchmark for LLM Memory Retrieval
במאמר זה נציגים נקודת ציון חדשה לבדיקת זיכרון LLM, המדדת דיוק החיפוש, ולא איכות התשובה. המאמר כולל גם תיאור של Tenure, פרוקסי לזיכרון-מאורגן, שמספק תכונות נוספות לבדיקה.
תקציר מקורי באנגליתarXiv:2605.11325v4 Announce Type: replace-cross Abstract: Current LLM memory benchmarks evaluate answer quality rather than retrieval accuracy. Consequently, a system that dumps its entire belief store can achieve perfect recall and mask severe precision failures. We show this evaluation gap persists across multiple embedding models where similarity-based retrieval over domain-specific corpora inherently struggles to isolate target beliefs from semantically proximate ones. Furthermore, multi-turn topic drift compounds this retrieval noise while driving up latency and operational costs. To decouple retrieval quality from generative performance, we introduce PrecisionMemBench, an 89-case benchmark measuring precision, noise isolation, session latency, and belief mutability. We also present T
קרא במקור המקורי