כתבה
arXiv cs.LG ·
שיפור ביצועים של אינפרנס מקומי
One Simple Trick for Improving the Performance of Energy-Limited Local Inference and Training
חוקרים מצאו דרך לשפר ביצועים של אינפרנס מקומי על ידי חלוקת העומס לחלקים קטנים. השיטה מוכיחה עצמה על מערכות כמו DGX Spark ושרתי מולטי-ג'יפי.
תקציר מקורי באנגליתarXiv:2609.11936v1 Announce Type: cross Abstract: Energy supply and heat dissipation are two of the main challenges with modern GPU deployments. While typically discussed in the context of new datacenter constructions, the same constraints also apply to small form-factor consumer devices, such as the DGX spark. In workloads characterized by alternating compute-intensive tasks such as matmuls with memory-bound operations such as norms or cross-entropy, the compute-intensive parts might hit power and/or thermal limits and start throttling. In this short paper, we show that chunking the workload into smaller parts that alternate compute and memory in higher frequencies, these power and temperature spikes can be smoothed out, preventing throttling and resulting in considerably faster wall-cloc
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית