כתבה
arXiv cs.CL ·
חידושי עבודה רציפה של מודלי שפה: חישוב עדפי עומק
Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
אנו מציגים חידושי עבודה רציפה של מודלי שפה, המאפשרים חישוב עדפי עומק יעיל. השיטה, הנקראת Continuous Depth Batching (CDB), מאפשרת למודלים לבצע חישובים עדפיים עומק, כאשר המודלים עובדים על טקסטים שונים. CDB משתמשת בטכנולוגיית חישוב רציפה, המאפשרת למודלים לבצע חישובים עדפיים עומק, כאשר המודלים עובדים על טקסטים שונים.
תקציר מקורי באנגליתarXiv:2608.09444v2 Announce Type: replace-cross Abstract: A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops cannot share a uniform forward pass and therefore cannot be handled by standard batching systems such as vLLM. The practical value of depth-adaptive inference thus hinges on whether batching can be made efficient. We introduce the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps. Our method dynamically schedules looped and non-looped parts of the architecture, manages looped KV-caching, and predict
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית