כתבה
arXiv cs.AI ·
מאמצי עיבוד: פיצול הדרך הראשונית מהעיבוד המשלים
Decode-Branch Transformers: Decoupling the Primary Prefill Path from Additional Decode Computation
מאמצי עיבוד: פיצול הדרך הראשונית מהעיבוד המשלים. פיתוח חדש של מודלי לשון גדולים, שמאפשר עיבוד נוסף ללא הגדלת העלויות.
תקציר מקורי באנגליתarXiv:2608.12385v3 Announce Type: replace Abstract: As large language models serve ever more requests, cumulative inference cost is growing relative to the one-time cost of training. In typical serving, prompt prefill runs in parallel and is compute-bound, whereas autoregressive decode is sequential and memory-traffic-bound. Conventional width or depth scaling raises both costs together, since every added layer is evaluated in both phases and enlarges the weights read at each decode step. We instead ask whether additional learned computation can be allocated to continuation prediction while preserving prompt-wide primary computation and a single KV cache. We realize this with the Decode-Branch Transformer. Its primary path alone processes the prompt and writes the KV cache; the decode bran
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית