יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

QK-Wanda: גזירה לא מובנית

QK-Wanda: Coupling Queries and Keys for Unstructured Pruning
QK-Wanda היא שיטה חדשה לגזירה של מודלי שפה גדולים. היא משפרת את שיטת Wanda על ידי ציון משקלים של שאילתות ומפתחות. השיטה נבדקה על מודלים כמו Llama ו-Qwen2.5.
תקציר מקורי באנגליתarXiv:2610.01554v1 Announce Type: new Abstract: Wanda (Sun et al., 2024) prunes large language models by scoring weights independently within each linear projection, although queries and keys interact through dot products. We introduce QK-Wanda, which scores query and key weights by their individual deletion costs under an unmasked pre-RoPE reconstruction objective. It augments Wanda scores with information from the opposite projection (keys for query weights, and queries for key weights), allowing both projections to share a pruning budget. Its closed-form scores require no gradients or weight updates; full pruning takes 1.3% longer than Wanda on A100 and 3.1% longer on H200 with the calibration used in our main experiments. We evaluate QK-only pruning across 15 models from TinyLlama, Lla
קרא במקור המקורי