כתבה
arXiv cs.AI ·
ThinQuant: רציונלי סקלבלי ללמידת הסתברות להקטנת ערכים ופעילות של LLMs
ThinQuant: Scalable Rotation Learning for Weight and Activation Quantization of LLMs
ThinQuant מציע רציונלי סקלבלי להקטנת ערכים ופעילות של LLMs. המאמר מציג פתרון חדש ללמידת הסתברות שמקטין את הזמן והמשאבים הדרושים להקטנת ערכים ופעילות של LLMs. ThinQuant משתמש בטכניקה של רציונלי סקלבלי כדי לקטין את הזמן והמשאבים הדרושים להקטנת ערכים ופעילות של LLMs. המאמר מציג תוצאות של ThinQuant שהצליחו להקטין את הזמן והמשאבים הדרושים להקטנת ערכים ופעילות של LLMs בכ-75%.
תקציר מקורי באנגליתarXiv:2609.36120v1 Announce Type: cross Abstract: Learned rotations play an important role in enabling low-bit weight and activation quantization of large language models by smoothing outliers in the activation distribution. State-of-the-art approaches include gradient-based procedures such as SpinQuant and computationally friendlier gradient-free approaches such as DartQuant, but both remain hard to scale to the largest architectures. To address the computational bottlenecks in gradient-free rotation learning, we introduce two ideas for efficiency, (i) a data selection procedure which reduces the required number of calibration data points, and (ii) an exact reduction of the associated optimization on this reduced calibration set. Our data selection procedure exploits the geometric structu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית