יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

LongSpark: דקודינג ספקולטיבי יעיל עם עלות קבועה לדראפטר פאראלל

LongSpark: Efficient speculative decoding with a fixed-cost parallel drafter
LongSpark מאפשר דקודינג ספקולטיבי יעיל עם עלות קבועה לדראפטר פאראלל. זה יעיל יותר מדראפטרים קיימים ומציע יתרון בעליות קונטקסט.
תקציר מקורי באנגליתarXiv:2609.37029v1 Announce Type: cross Abstract: Speculative decoding accelerates autoregressive inference by verifying multiple draft tokens in a single target forward pass. However, as the context grows, existing state-of-the-art drafters become increasingly expensive, eroding the very efficiency advantage they are designed to provide. We argue that this scaling is unnecessary. A standalone language model must grow with its prefix because it is solely responsible for every token it produces. A drafter, by contrast, only proposes candidates; the target catches and corrects every error before any token is committed. The drafter's decoding cost can therefore be made entirely independent of the prefix length. We introduce LongSpark, a block-diffusion drafter that achieves this by extracting
קרא במקור המקורי