יום רביעי, 7 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

התאמה, אז תיקון: תרגום-חינם של שני-שלבים לפיצוי נמוך-דרגה למודלי שפה גדולים מאוד

Align, Then Correct: Training-Free Two-Stage Low-Rank Compensation for Extremely Quantized Large Language Models
במאמר זה, המחברים מציגים פרקטיקה חדשה לפיצוי נמוך-דרגה למודלי שפה גדולים, המאפשרת תרגום-חינם ללא צורך באימון. הם מציעים שני שלבים: התאמה, שבו הם רושמים את המודל המלא-דיונם, ותיקון, שבו הם תוקנים את המודל המקומם. הם מדגימים את הפרקטיקה על שני מודלי שפה גדולים, Qwen3-8B ו-Qwen3-4B, ומציגים תוצאות טובות יותר מאשר המודלים המקוריים.
תקציר מקורי באנגליתarXiv:2610.08164v1 Announce Type: new Abstract: Low-rank quantization error compensation (LQEC) recovers the accuracy lost under aggressive weight quantization by attaching a closed-form rank-$r$ adapter beside each frozen quantized weight, without any training. We show that existing compensators are limited by two shared simplifications. They calibrate symmetrically, evaluating the full-precision and compensated weights on the same activation, which yields a compensation target that is inherently high-rank -- so a fixed rank budget captures only a small fraction of it. And they minimize only the second-order term of the loss, although the compensated model is not stationary: a first-order descent direction larger than the applied compensation itself remains in every layer, and no reconstr
קרא במקור המקורי