יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

הפרדת עיכוב-פלט: מערכת חינם לבקר והגבלות כלליות

Decode-Latency Feedback Prefill: A Model-Free Controller and Its Generalization Limits
מערכת חינם לבקר עיכוב-פלט, המפחיתה את עיכוב הפלט במסגרת הסתברותית. המערכת נבחנה על ידי טכנולוגיית NVIDIA והציגה תוצאות טובות. עם זאת, המערכת לא הצליחה לפעול כראוי עם מודלים גדולים יותר.
תקציר מקורי באנגליתarXiv:2609.38386v1 Announce Type: new Abstract: Concurrent autoregressive inference creates a fundamental interference problem: prefilling a newly arrived long prompt can delay tokens for requests that are already decoding. Fixed prefill chunks reduce this interference, but the best chunk size depends on the model, hardware, load, and latency objective. We introduce Decode-Latency Feedback Prefill (DLFP), a model-free controller that changes only prefill work that overlaps active decodes. After a guarded scheduling cycle, DLFP uses the observed interval as proportional feedback to resize the next prefill chunk; isolated prefills remain unrestricted. We implement DLFP in vLLM and evaluate it with open-loop Poisson arrivals, exact token accounting, raw request traces, and NVIDIA telemetry. O
קרא במקור המקורי