כתבה
arXiv cs.LG ·
הסקה עמוקה מותאמת עומק למודלים שפה מחוגרים
Depth-adaptive Inference of Looped Language Models via Continuous Depth Batching
חוקרים פיתחו שיטה חדשה להסקה עמוקה מותאמת עומק עבור מודלים שפה מחוגרים. השיטה, Continuous Depth Batching, מאפשרת למודלים להשתמש בפחות חישוב עבור טוקנים 'קלים' ויותר עבור 'קשים'.
תקציר מקורי באנגליתarXiv:2608.09444v2 Announce Type: replace Abstract: A main promise of looped language models is depth-adaptive inference. By looping a block of shared layers a variable number of times, the model can use less compute for "easy" tokens and more for "hard" ones. However, tokens with different numbers of loops cannot share a uniform forward pass and therefore cannot be handled by standard batching systems such as vLLM. The practical value of depth-adaptive inference thus hinges on whether batching can be made efficient. We introduce the first efficient method for depth-adaptive looped LMs via continuous depth batching (CDB), which forms new batches between loop steps. Our method dynamically schedules looped and non-looped parts of the architecture, manages looped KV-caching, and predicts whic
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית