כתבה
arXiv cs.AI ·
DynaTE: האצת LLMs דיפוזיוניים
DynaTE: Accelerating Diffusion LLMs via Dynamic Token Execution
DynaTE היא ארכיטקטורה חדשה שמאיצה LLMs דיפוזיוניים. היא משתמשת בקיזוז טוקנים בעלי תועלת נמוכה ומיטוב זרימת הנתונים. DynaTE משיגה האצה של 2.05-2.78 פעמים ויעילות אנרגיה גבוהה יותר מאשר מאיצים אחרים.
תקציר מקורי באנגליתarXiv:2610.11284v1 Announce Type: cross Abstract: Diffusion-based LLMs (dLLMs) have recently emerged as a promising alternative to autoregressive (AR) LLMs by enabling bidirectional parallel refinement, alleviating the sequential decoding bottleneck of AR generation. However, their parallel iterative refinement mismatches AR accelerators optimized for sequential decoding and their discrete token generation differs from DiT accelerators designed for continuous denoising. Recent dLLM accelerators have explored workload-specific optimizations to reduce vocabulary processing overhead and redundant computation across denoising iterations. However, these approaches retain all tokens in parallel execution, despite varying token refinement utility and execution requirements. This paper presents Dy
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית