כתבה
arXiv cs.LG ·
רפרמטריזציה של סופטמקס להקטנת ראשי-תוצאה
Softmax Reparameterization for Output-Head Quantization
במאמר זה, המחברים מציגים טכניקה חדשה להקטנת ראשי-תוצאה של דגלי שפה. הם מציעים רפרמטריזציה של סופטמקס, שהיא שיטה שמחפשת ראשי-תוצאה פונקציונלית שווה למקורי, לפני הקטנה. השיטה מחסרת פער של סקלר מרובע-משותף של ווקבולרי-שור משורה-משותף, מכל רשומה-תוצאה, ובוחרת את הקואפיצנט על ידי תקינה KL. השיטה נותנת תוצאות-פרדיקציה מדויקות לחלוטין לראשי-תוצאה-לינארי-סופטמקס, ואינה דורשת רטריינינג מחדש של המקורי-מפענח.
תקציר מקורי באנגליתarXiv:2609.31291v2 Announce Type: new Abstract: Large vocabularies make output heads a substantial inference cost in small language models. We introduce softmax reparameterization, a post-training method that searches over functionally equivalent output heads before quantization. The method subtracts a scalar multiple of the vocabulary-row mean from every output row and selects the coefficient by validation KL. For linear-softmax heads, these shifts preserve full-precision predictions exactly and require no decoder retraining; a rank-one correction extends the construction to nonlinear logit paths. Across seven output heads and three quantizers, W4 gains are largest where baseline quantization substantially distorts predictions: test KL falls by 93% on XGLM under RTN and by 73--77% on Phi,
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית