כתבה
arXiv cs.CL ·
Prox: Training-Free FFN Activation Sparsity via Approximate Intermediate-Channel Salience in LLMs
תקציר מקורי באנגליתarXiv:2607.27591v1 Announce Type: cross Abstract: Feed-forward networks (FFNs) dominate memory traffic and computation in large language model (LLM) inference, making them a primary target for activation sparsification. However, existing training-free methods suffer substantial model-quality degradation at high sparsity due to limitations in their channel-selection strategies. We observe that the SwiGLU intermediate state provides a highly effective channel-selection signal, but obtaining it requires costly dense computation. To address this, we present \emph{Prox}, a two-stage training-free framework for sparse SwiGLU FFNs. Prox hinges on the key insight: sparse execution requires only the channel mask induced by the intermediate state, which can be constructed from the magnitude ranking
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית