כתבה
arXiv cs.LG ·
LRCC: דחיסת דירוג נמוך עם חישוב מותנה
LRCC: Generalizing Low-Rank Compression with Conditional Computation
LRCC היא שיטה חדשה לדחיסת מודלי שפה מוקדמים. היא משתמשת בחישוב מותנה על פי טוקנים, ומשפרת את ביצועי המודלים Llama ו-Qwen.
תקציר מקורי באנגליתarXiv:2610.08858v1 Announce Type: cross Abstract: Low-rank compression reduces the cost of pretrained language models by replacing linear transformations with low-rank factorizations. However, conventional methods use a fixed rank allocation during inference, assigning the same amount of compute regardless of the input token. We introduce Low-Rank Conditional Computation (LRCC), which adds token-dependent computation to pretrained models by training one lightweight router per Transformer block to select among a small set of nested low-rank paths. During training, the low-rank factors remain frozen, and only the routers are optimized. We evaluate LRCC on Llama and Qwen models for language modeling and zero-shot downstream tasks. Within the same average active-parameter budget, LRCC improves
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית