כתבה
arXiv cs.CL ·
PELM: פתרון יעיל לביצועי LLM במכשירי קצה
PELM: Power Efficient On-Device LLM Inference with Speculative Decoding and Dynamic Voltage Frequency Scaling
PELM היא פתרון יעיל לביצועי LLM במכשירי קצה, המשתמשת בטכניקות של קודינג ספקולטיבי ושינוי תדר ותדירות דינמי. הפתרון נועד לקצר את זמן הביצוע ולהפחית את צריכת האנרגיה של LLM במכשירי קצה.
תקציר מקורי באנגליתarXiv:2609.09662v1 Announce Type: cross Abstract: Deploying Large Language Models (LLMs) directly on mobile platforms at the edge is gaining traction due to a myriad of benefits, such as increased privacy, personalization, and reduced latency. However, LLMs have heavy computational requirements, which are difficult for resource-constrained mobile and edge platforms to fulfill. In addition to limited compute resources, mobile and edge systems often have a compact form factor and lack physical mechanisms to dissipate heat generated from high processor usage rates (e.g., fans) to prevent throttling and reduced processing power, which LLMs can easily cause. To mitigate these effects, prior works have proposed various power governing strategies, such as dynamic voltage and frequency scaling (DV
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית