יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

TopK-Guided: רזונות מותאמים למודלי LLM

TopK-Guided: Adaptive, Budget-Aware Activation Sparsity for Efficient LLM Inference
TopK-Guided הוא שיטה חדשה לרזונות מותאמים למודלי LLM. השיטה משלבת רזונות ברמת הטוקן עם חלוקת תקציב רגישה ברמת הבלוק. היא משפרת את הביצועים של מודלי Llama-2 ו-Llama-3.
תקציר מקורי באנגליתarXiv:2610.01763v1 Announce Type: new Abstract: Activation sparsity speeds up large language model (LLM) inference by setting unimportant activations to zero so that the corresponding computations can be skipped. Existing training-free methods, however, make different trade-offs: threshold-based methods such as TEAL adapt the sparsity level to each token but do not tightly control the realised sparsity, while TopK-based methods such as WINA enforce a fixed sparsity level but use the same sparsity budget for every token. Both also apply the same budget across transformer blocks, despite large differences in block sensitivity. We introduce TopK-Guided, a training-free method that addresses both limitations by combining bounded token-level sparsity adaptation with sensitivity-aware block-leve
קרא במקור המקורי