כתבה
arXiv cs.AI ·
זיכרון קבוע במערכת יישום של LLM: מה זה עולה, מה זה קונה, ומתי יש לדעת
Persistent Memory in Multi-Agent LLM Inference: What It Costs, What It Buys, and When You Can Tell
במאמר זה נחקרה דרך לפחתן את זיכרון ה-KV במערכת יישום של LLM. התוצאות הראו שהשיטה פועלת, אך השיפור בזיכרון ה-KV אינו ניכר באופטימיזציה של הדיוק.
תקציר מקורי באנגליתarXiv:2610.07782v1 Announce Type: new Abstract: Decomposing long-context inference across cooperating agents bounds the active KV cache per call rather than total evidence, which matters when KV-cache memory binds. Many such systems add a persistent tier storing and recalling reasoning traces, usually validated by an ablation reporting an accuracy gain. We measure both on one three-tier agent architecture. Decomposition delivers: peak KV working set of 14.3 MiB per query against 35.5 and 35.3 MiB for single-pass and retrieval-augmented baselines. The persistent tier does not: across eight controlled dataset pairs at n=100 per arm it costs +0.368 MiB [+0.167, +0.590] of peak cache and produces no detectable accuracy change (+0.015, 95% CI [-0.011, +0.046]). We argue the null is structural:
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית