יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

כשההפלה הפאשיסטית לא עובדת: חשוב מחדש את תחליפי הזיכרון לשימוש מראש של LLM

When Fancy Eviction Fails: Rethinking Cache Replacement For LLM Prefix Reuse
אנחנו חוקרים את תחליפי הזיכרון לשימוש מראש של LLM ומצאנו שפוליטיקות מתקדמות לא יותר משפרות את LRU. ניתן לשמור רצינות כאשר נותנים פרישה מהירה לפרפקסים של פעם אחת, פרישה מודעת לחישוב לפרפקסים יקרים וגרנולריות של פרישה תלויות בקיבולת.
תקציר מקורי באנגליתarXiv:2609.28870v2 Announce Type: replace-cross Abstract: Long-running LLM applications repeatedly send growing context, making prefix caching critical for reducing prefill cost. Yet prefix-cache behavior under agentic workloads remains poorly understood. We study production traces from two companies and evaluate 14 eviction algorithms across HBM-constrained and large memory-pool settings. Despite a large gap to Belady, sophisticated policies designed for traditional caches provide little benefit over LRU. The reason is structural: prefix reuse is dominated by the regular pacing of active sessions, making recency unusually predictive. Prefix caching nevertheless introduces new challenges, including heavy-tailed session footprints and highly variable miss costs as attention computation grow
קרא במקור המקורי