כתבה
arXiv cs.AI ·
Compressing What Matters: Neuron Importance Meets Data-Aware Low Rank Approximation for Language Model Compression
תקציר מקורי באנגליתarXiv:2607.18284v1 Announce Type: cross Abstract: To excel at their domain large language models are comprised of billions of parameters. Yet this comes at the cost of huge memory requirements restricting their applicability in resource-constrained environments. To address the problem of neural network (NN) compression Singular Value Decomposition (SVD) has played a key role as a fundamental component for matrix compression through decomposition. To minimize compression error and to maximize the efficacy of the compressed model on the downstream tasks previous works focused on low-rank approximation of the NN's weight matrices either from the perspective of parameter importance or per-layer functional equivalence. While previous works studied the aforementioned perspectives in isolation in
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית