כתבה
arXiv cs.AI ·
LongSpark: דקודינג ספקולטיבי יעיל עם עלות קבועה לדראפטר פאראלל
LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter
LongSpark מאפשר דקודינג ספקולטיבי יעיל עם עלות קבועה לדראפטר פאראלל. זה יעיל יותר מדראפטרים קיימים ומציע יתרון בעליות קונטקסט.
תקציר מקורי באנגליתarXiv:2609.37029v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is unnecessary. A standalone language model must grow with its prefix because it is solely responsible for every token it produces. A drafter, by contrast, only proposes candidates; the target catches and corrects every error before any token is committed. The drafter's decoding cost can therefore be made entirely independent of the prefix length. We introduce LongSpark, a block-diffusion drafter that achieves this by extracting
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית