כתבה
arXiv cs.LG ·
LLMET: הפקת תצוגה חסכונית באנרגיה של LLM על ידי שימוש בזיכרון M3D
LLMET: Enabling Cross-Layer Evaluation of Emerging M3D Memories for Energy-Efficient LLM Serving
LLMET הוא פלטפורמת השקעה שמטרתה לצמצם את צריכת האנרגיה של LLM על ידי שימוש בזיכרון M3D. הפלטפורמה נבחנה על ידי צוות מדענים שהשוו את צריכת האנרגיה של LLM עם ובלי שימוש בזיכרון M3D. התוצאות הראו ירידה ניכרת בצריכת האנרגיה של LLM.
תקציר מקורי באנגליתarXiv:2607.26491v1 Announce Type: cross Abstract: The energy consumption of Large Language Model (LLM) serving is becoming a major system challenge as deployment scales, driven by hardware power and thermal constraints and rising electricity costs. A key contributor to chip energy dissipation is data movement between limited on-chip cache and off-chip High Bandwidth Memory (HBM). Meanwhile, emerging memory technologies such as monolithic 3D (M3D) integration of cache memories at the Back-End-Of-Line (BEOL) of logic chips enable larger and denser on-chip memories, creating new opportunities to reduce costly off-chip traffic. However, it remains unclear whether continuously scaling on-chip memory using emerging technologies can effectively improve the energy efficiency of LLM serving. To add
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית