יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

קיצוץ זרם שאריות תוך התחשבות בפלט

Output-aware Residual Stream Pruning for Large Language Models
שיטה חדשה לקיצוץ זרם שאריות במודלים גדולים של שפה, המתחשבת ברגישות של שכבות היצוא. השיטה משפרת את התוצאות ואת היעילות בהשוואה לשיטות קודמות.
תקציר מקורי באנגליתarXiv:2609.35579v2 Announce Type: replace Abstract: Residual stream pruning methods reduce inference cost by shrinking the model's hidden dimension, but existing approaches typically choose these dimensions by minimizing activation reconstruction error. This criterion implicitly treats all perturbation directions as equally important, ignoring the sensitivity of downstream layers. We introduce a sensitivity-aware approach to residual-stream pruning that directly accounts for this direction-dependent sensitivity. Using a second-order approximation to the output KL divergence, we characterize the effect of a residual-stream perturbation through both its activation covariance and the local sensitivity of the model output. The resulting subspace selection objective couples these two quantities
קרא במקור המקורי