יום שישי, 9 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

SparseDecoding: פרוש ניהולי לפרוש ניהולי למודלי LLM

SparseDecoding: Decoding-Aware Pruning for Accurate and Efficient LLM Inference
פרוש ניהולי לפרוש ניהולי למודלי LLM: פרוש ניהולי לפרוש ניהולי למודלי LLM
תקציר מקורי באנגליתarXiv:2610.12327v1 Announce Type: new Abstract: The memory-bound nature of the decoding stage of large language model (LLM) inference incurs significant latency. Layer-wise training-free network pruning approaches guided by the Hessian have been a prominent solution to this problem, as pruning reduces the number of nonzero parameters read from memory during decoding. Nevertheless, typical methods in this line compute the Hessian using pre-collected natural sequences, whereas the model is fed self-generated tokens during decoding, creating a distribution shift between the two sequences. The Hessian calculated on the natural sequence is different from that calculated on the generated sequence. We observe that this discrepancy causes the activation distribution during generation to deviate fr
קרא במקור המקורי