כתבה
arXiv cs.LG ·
SQS: דחיסה בספקטרום Bayesiano של DNN דרך ספארסיות מצומצמות וקוונטיזציה נמוכה
SQS: Bayesian DNN Compression through Sparse Quantized Sub-distributions
אורחות דחיסה Bayesiano של DNN דרך ספארסיות מצומצמות וקוונטיזציה נמוכה. החידוש נעשה על ידי ResNet, BERT-base, Llama3.2 ו-Qwen2.5.
תקציר מקורי באנגליתarXiv:2510.08999v2 Announce Type: replace Abstract: Compressing large-scale neural networks is essential for deploying models on resource-constrained devices. Most existing methods adopt weight pruning or low-bit quantization individually, often resulting in suboptimal compression rates to preserve acceptable performance drops. We introduce a unified framework for simultaneous pruning and low-bit quantization via Bayesian variational learning (\method), which achieves higher compression rates than prior baselines while maintaining comparable performance. The key idea is to employ a spike-and-slab prior to induce sparsity and model quantized weights using Gaussian Mixture Models (GMMs) to enable low-bit precision. Due to the intractability of the objective involving spike-and-slab priors wi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית