כתבה
arXiv cs.LG ·
Activation Sparsity with Weight Approximation for Faster LLM Decoding on Offloaded Weights
תקציר מקורי באנגליתarXiv:2610.02598v1 Announce Type: new Abstract: Deploying LLMs on consumer-grade GPUs with insufficient memory to hold their weights can result in prohibitively slow inference, because decoding repeatedly transfers offloaded weights from system RAM or flash storage into GPU at much lower bandwidth than local GPU-memory access. Activation sparsity reduces these transfers by skipping weights associated with zero or near-zero activations. However, as more activation contributions are omitted, model quality eventually degrades rapidly, indicating that weights associated with small-magnitude activations collectively influence model quality sharply. In this work, we improve the trade-off between model quality and decoding performance when exploiting activation sparsity. Our key idea is to replac
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית