יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

SpecFold: קיפול ערוצים מרובים לפיענוח מהיר

SpecFold: Folding Multi-Branch Redundancy for Faster Speculative Decoding in Diffusion Language Models
SpecFold הוא אלגוריתם שמאיץ את דגמי השפה הדיפוזיים. הוא עובד עם מודלים כמו LLaMA ומשתמש בטכנולוגיות כמו LangChain. SpecFold מקנה ביצועים מהירים יותר בפיענוח טקסט.
תקציר מקורי באנגליתarXiv:2610.04875v2 Announce Type: replace Abstract: Diffusion large language models (DLLMs) generate text through iterative block denoising, and multi-branch speculative decoding accelerates this process by verifying a main branch together with multiple draft branches in a single forward pass. While prior DLLM acceleration methods primarily exploit temporal redundancy across denoising steps, we identify a complementary redundancy axis within each speculative verification step: multi-branch computational redundancy. During speculative verification, draft branches inherit most tokens from their parents while unmasking a small set of additional positions, causing large portions of hidden states to remain highly similar across branches. We propose SpecFold, an algorithm-system co-design that e
קרא במקור המקורי