יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

קראת חישוב מבוזרת ל-LLM

Toward Sustainable Distributed LLM Inference: A Systems Synthesis and Research Agenda for an Energy-, Carbon-, and Cache-Aware llm-d Control Plane
החידוש עוסק בפיתוח קראת חישוב מבוזרת למודלי שפה גדולים. המחקר בוחן את השפעת גורמים שונים על צריכת האנרגיה ופליטת הפחמן של מודלים אלה, ומציע פתרון לניהול חישוב יעיל ובר קיימא.
תקציר מקורי באנגליתarXiv:2609.05565v1 Announce Type: cross Abstract: Large language model (LLM) sustainability is increasingly a serving-systems problem, not only a training problem. In production, energy and carbon impact depend on more than model size: workload shape, batching, key-value (KV) cache reuse, prefill/decode placement, model and accelerator choice, power state, geographic carbon intensity, and service-level objectives (SLOs) all matter. Recent systems papers study many of these factors separately. This paper connects those results and asks a practical engineering question: what do they imply when the decision point is a distributed inference control plane such as llm-d? The contribution here is synthesis, not a new set of benchmark results. Reported performance, energy, carbon, and cost improve
קרא במקור המקורי