כתבה
arXiv cs.AI ·
BitNest: פיתוח לתאוריה של פענוח ספקולטיבי
BitNest: Bit-Nested Speculative Decoding for Memory-Efficient LLM Inference Acceleration
BitNest הוא פיתוח חדש שמאפשר פענוח ספקולטיבי מהיר יותר עבור מודלים של LLM, כולל LLaMA. הוא משתמש בשיטה חדשה ליצירת גרסה נמוכת דיוק של המודל, ואז משפר אותה לגרסה בעלת דיוק גבוה יותר. זה מאפשר ל-BitNest לחסוך זיכרון ולהאיץ את תהליך הפענוח.
תקציר מקורי באנגליתarXiv:2610.02800v2 Announce Type: replace Abstract: Speculative decoding accelerates autoregressive generation by using a lightweight draft to propose multiple tokens for parallel verification. However, existing methods often require an additional draft model or weight representation, introducing non-negligible memory overhead on resource-constrained devices. Self-speculative approaches reduce this overhead, yet still face trade-offs between draft quality, target quality, and storage efficiency. We propose BitNest, a bit-nested speculative decoding framework that embeds a low-precision draft directly into the higher-precision target representation. Instead of deriving a draft from a predefined target, BitNest first constructs a strong low-precision base and then recovers the higher-precisi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית