יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

ConQuR: פענוח פעילות מאורגן על ידי סיבוב אופטימלי למודלי LLM

ConQuR: Corner Aligned Activation Quantization via Optimized Rotations for LLMs
ConQuR הוא שיטה חדשה לפענוח פעילות של מודלי LLM, המשתמשת בסיבוב אופטימלי כדי להפחית את הטעות בפענוח. השיטה נבחנה במודלי Llama-2 ו-Llama-3 והוכיחה עצמה כיעילה ויעילה. ConQuR יכולה לשפר את הביצועים של מודלי LLM ולהפחית את העלויות שלהם.
תקציר מקורי באנגליתarXiv:2605.10793v2 Announce Type: replace Abstract: Large language models (LLMs) are costly to deploy due to their large memory footprint and high inference cost. Weight-activation quantization can reduce these costs, but low-bit activation quantization remains difficult because activation outliers induce large quantization error. Recent rotation-based methods address this by applying orthogonal transformations that redistribute activation magnitude across dimensions, but existing approaches either require expensive end-to-end rotation training or rely on stored activation corpora, introducing significant compute or storage overhead. We propose a lightweight post-training rotation calibration method for LLM activation quantization. Our method learns orthogonal rotations that align normaliz
קרא במקור המקורי