כתבה
arXiv cs.LG ·
SSR: Sparse Segment Reduction for Ternary GEMM Acceleration
תקציר מקורי באנגליתarXiv:2610.08403v1 Announce Type: new Abstract: Large Language Models (LLMs) require substantial computational resources, limiting their deployment on resource-constrained hardware. Ternary LLMs mitigate these demands through weight quantization via ternary values, achieving significant compression often with 50-90% sparsity. However, existing approaches have limitations: methods optimized for ternary weights, such as BitNet, redundant segment reduction (RSR), and its improved version RSR++, do not exploit sparsity structures, while conventional sparse formats neglect ternary characteristics, foregoing dual optimization opportunities. In this paper, we introduce Sparse Segment Reduction (SSR), a ternary matrix multiplication method designed to accelerate the inference of ternary LLMs and g
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית