כתבה
arXiv cs.AI ·
Reconstructing the Right Episode: Evaluating Interleaved Conversational Memory Beyond Long Context
תקציר מקורי באנגליתarXiv:2608.25655v2 Announce Type: replace-cross Abstract: Conversations with chat assistants increasingly span many topics in a single long-running thread, challenging memory systems. Existing long-context and memory benchmarks often expose session or topic boundaries, or probe direct personal-memory questions. These settings understate a harder assistant-memory regime: a flat mixed-topic thread where the system must infer which earlier episode makes a later task decision valid. We introduce SCALE-QA, a constraint-grounded task QA benchmark for flat unsegmented threads targeting episode integrity failure. The dataset contains 3,000 audited questions across 10 domains, uses deterministic four-way multiple-choice grading, and includes a deterministic runtime builder; experiments use all 3,00
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית