יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

Hardware-Native Joint Sparse-Quantization for Trillion-Scale Mixture-of-Experts

תקציר מקורי באנגליתarXiv:2610.02241v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) architectures allow frontier language models to scale to trillions of parameters, but their deployment is constrained by massive memory footprints and memory-bandwidth limitations. Although modern accelerators provide Sparse Tensor Cores (SpTCs) that reduce weight storage and increase throughput through low-precision semi-structured sparsity, exploiting them for MoEs remains challenging because of substantial model-quality degradation and the lack of grouped sparse GEMM primitives. We present an end-to-end hardware-software co-design framework that compresses expert weights into hardware-native, low-precision sparse representations and accelerates their execution on SpTCs. Algorithmically, our framework relaxes disc
קרא במקור המקורי