כתבה
arXiv cs.LG ·
זיכרון קבוע במערכת LLM: מה זה עולה, מה זה קונה, ואיזה זמן יש לדעת
Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell
במאמר זה, נחקר השימוש בזיכרון קבוע במערכת LLM, ונבחן האם הוא עולה ומה זה קונה. התוצאות היו שלא נמצאה תועלת בשימוש בזיכרון קבוע.
תקציר מקורי באנגליתarXiv:2610.07782v1 Announce Type: cross Abstract: Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain. We measure both on one three-tier agent architecture. Decomposition delivers: peak KV working set of 14.3 MiB per query against 35.5 and 35.3 MiB for single-pass and retrieval-augmented baselines. The persistent tier does not: across eight controlled dataset pairs at n=100 per arm it costs +0.368 MiB [+0.167, +0.590] of peak cache and produces no detectable accuracy change (+0.015, 95% CI [-0.011, +0.046]). We argue the null is structural
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית