יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

בדיקת זיכרון ארוך-טווח: שיפוט מחדש, שונות בקוראים ובדיקות שליליות

Auditing Long-Term Memory Evaluation: Repeated Judging, Reader Variation, and Negative Controls
בדיקת זיכרון ארוך-טווח: שיפוט מחדש, שונות בקוראים ובדיקות שליליות. המחקר חוקר את תהליך הבדיקה של זיכרון ארוך-טווח ומצא כי הוא כולל שיפוט מחדש, שונות בקוראים ובדיקות שליליות. המחקר גם מצא כי הזיכרון הארוך-טווח לא ניתן לבדיקה באופן יעיל.
תקציר מקורי באנגליתarXiv:2609.38021v2 Announce Type: replace Abstract: This report audits evaluation of a long-term-memory retrieval chain on the 500 LongMemEval-S development questions. Its strongest historical reader lane scores 479 and 475 under an adapted GPT-4o rubric; re-judging the same pass-1 answers changes three labels and yields 478. Fixed-answer knowledge-update re-scoring gives 70/72 under the upstream template and 69/72 under the modified template. Reader lanes span 93 to 479 on fixed packets; paired tests between the two strongest historical lanes establish neither superiority nor equivalence. A different-family reader, configured without client tools or operator files, scores 474, 1.0 percentage point below the headline pass (paired 95% interval [-3.0,+1.0]). Live reader request bodies were n
קרא במקור המקורי