יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

SparseEngine: מנוע השקעה ספארס

SparseEngine: Sparse-First Inference Engine
SparseEngine הוא מנוע השקעה ספארס שנועד לטפל במערכות LLM עם היסטוריות אינטראקציה ארוכות. הוא מספק עד 10x הגברת תעבורה עם השקעה ועד 2.5x עלייה במהירות דקודינג.
תקציר מקורי באנגליתarXiv:2609.39068v1 Announce Type: new Abstract: Long-context LLM agents accumulate interaction histories that strain KV-cache memory and attention computation. Although sparse attention reduces these costs, heterogeneous cache representations and workflows hinder integration with existing inference engines, while prior sparse-serving abstractions support only specific layouts or workflows. We present SparseEngine, a ground-up, sparse-first inference engine whose shared lifecycle contract lets each method control its KV representation and computation while coordinating state transitions with common serving infrastructure. SparseEngine supports 15 methods across four categories and enables cross-request state management through Chain Cache, which resumes KV-eviction methods from retained his
קרא במקור המקורי