כתבה
arXiv cs.AI ·
אופטימיזציה של LLM Serving עם גודל קבוע של Prefill ו-Decode
LLM Serving Optimization with Variable Prefill and Decode Lengths
במאמר זה, המחברים חקרו את האופטימיזציה של LLM Serving תחת תקציב זיכרון KV-קבוע, כאשר הבקשות הן חסרות-סדר וגודלן של הפרטים (prefill) והתשובות (decode) הן שונות. הם הציגו פוליסה חדשה, Sorted-F, שמחפשת תרחישים ישימים באמצעות F-מטריקה שמאזן את קרדינליות התרחיש כנגד עלות הקודקוד. הם גם פיתחו תוכנית דינמית פשוטה לבעיה הסטטית, וגם גרפים חיפוש-מקומי, גרפים חכם, וגרפים LP-מנחים.
תקציר מקורי באנגליתarXiv:2508.06133v5 Announce Type: replace-cross Abstract: We study offline scheduling for large language model (LLM) serving under a fixed KV-cache memory budget, where requests have heterogeneous prompt (prefill) and response (decode) lengths. Given a backlog of requests available at time zero, the scheduler forms mixed prefill/decode batches over time to minimize total end-to-end latency. We show that heterogeneity in prompt lengths fundamentally changes the problem: minimizing total latency is NP-hard, and standard policies that prioritize short outputs or small total sequence sizes can have unbounded approximation ratios. We propose Sorted-F, which repeatedly selects feasible batches using an F-metric that balances batch cardinality against downstream decode cost. With exact batch sele
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית