יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

צינון-משוכלל של פיפליין להפרעה-LLM ב-GPU יחיד: שיטות, תרחישים והשפעות קשר

Budget-Aware Compression Pipeline for Single-GPU LLM Inference: Methods, Trade-offs, and Coupling Effects
מחקר חדש מציג שיטות להפחתת גודלן של מודלי שפה גדולים על גבי GPU יחיד. המחקר עוסק בפיתוח פיפליין צינון-משוכלל להפרעה-LLM.
תקציר מקורי באנגליתarXiv:2608.30076v2 Announce Type: replace Abstract: Single-GPU deployment of 70B-parameter language models on an NVIDIA GPU is constrained by device memory, long-context throughput, and engineering integration cost. We cast single-GPU inference as a budget-aware design problem over these three axes and study how pruning, quantization, and KV-cache compression interact under realistic execution. Controlled ablations show that layer-wise pruning makes weight quantization more robust. KV-cache sparsification complements INT8 KV quantization by reducing memory without hurting decoding speed, while static vector quantizers often conflict with dynamic caching. Guided by these coupling results and explicit budget tracking, we assembled a practical pipeline and compressed a 70B model to about 33 G
קרא במקור המקורי