כתבה
arXiv cs.LG ·
פיענוח מונחה: אינפרנס מהיר פוגש תגבור
Mentored Decoding: Faster Inference meets Boosting
פיענוח מונחה הוא שיטה להאצת אינפרנס של מודלי שפה אוטורגרסיביים. המחקר מציג גישה פורמלית לפיענוח מונחה, המאפשרת לשפר את מהירות האינפרנס ואת איכות התוצאות.
תקציר מקורי באנגליתarXiv:2609.30474v1 Announce Type: new Abstract: Speculative decoding is a successful technique speeding up inference of a target autoregressive language model via a fast drafter model. Lossy speculative decoding allows a drift with respect to the target to further improve speed. Interestingly, it has been observed experimentally that the resulting model can $\textit{also}$ beat the target $\textit{quality-wise}$. Our paper formally proves how such a feat is possible with a formal approach to lossy speculative decoding called $\textit{mentored decoding}$. To get there, we connect inference to a celebrated ML training theory, $\textit{boosting}$, and proceed via the generalization of mentored decoding to the whole set of $f$-divergences. We uncover key properties of mentored decoding, among
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית