כתבה
arXiv cs.LG ·
אופטימיזציה מדויקת של זיכרון-זמן לשרתי מודלי שפה
Exact Memory-Time Optimization for Prefix-Cached Language Model Serving
חוקרים פיתחו שיטה חדשה לאופטימיזציה של שרתי מודלי שפה, המשתמשת בקידוד קדמי ואחסון זיכרון. השיטה, הנקראת Prefix-Certificate Retention, מאפשרת לשפר את יעילות השרתים ולחסוך זיכרון.
תקציר מקורי באנגליתarXiv:2610.02766v1 Announce Type: new Abstract: Retaining language-model prefix states trades recomputation against storage time. Optimizing each cached block independently can overcount savings: a resident block is usable only when the required preceding prefix is also available. We introduce Prefix-Certificate Retention (PCR), an exact finite-trace formulation for static, grouped, reset-on-access timeouts. Usable-prefix rewards become nodes whose prerequisites are timeout thresholds and preceding hit certificates. The resulting maximum-weight closure reduces to one minimum cut, with graph size linear in the number of block lookups and timeout choices. A breakpoint theorem extends the construction to all nonnegative timeouts without discretization error. We also derive a linear-time-in-gr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית