כתבה
arXiv cs.LG ·
SketchSSM: Write to the Full State, Read from a Compact Sketch
תקציר מקורי באנגליתarXiv:2609.33051v2 Announce Type: replace Abstract: Hybrid-attention models replace most softmax attention layers with linear attention, reducing KV-cache growth and enabling larger decode batches where recurrent-state access becomes a major bottleneck. ReplaySSM amortizes state updates by buffering keys and values, but each new query still requires a full-state read even though the state remains unchanged between state updates. We observe that low-rank state-weighted query approximation accurately preserves state-read outputs. Although future queries are unknown, the basis vectors used to approximate them can be fixed offline. Based on this observation, we introduce SketchSSM, which preserves full-state updates while approximating reads. At each state update, SketchSSM reads the full stat
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית