יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

OnlineQAT: שיטה חדשה לדחיסת מודלי שפה גדולים

OnlineQAT: On-Policy Distillation for Ultra-Low-Bit Large Language Models
OnlineQAT היא שיטה חדשה לדחיסת מודלי שפה גדולים, כגון Qwen3-1.7B, למצבי נמוכי עומק. השיטה משלבת אימון חוזר ודחיסה על מנת לשפר את הדיוק. תוצאות הניסויים מראות כי OnlineQAT משפרת את הדיוק ב-2.90 נקודות ב-W3A16 ו-0.44 נקודות ב-W2A16, לעומת שיטות אחרות.
תקציר מקורי באנגליתarXiv:2610.09346v1 Announce Type: new Abstract: Quantization-aware training (QAT) can recover much of the accuracy lost when large language models are compressed below four bits. Existing re- covery stages, however, are commonly optimized on fixed completions or teacher-generated answers, whereas the deployed quantized model condi- tions on prefixes generated by itself. Quantization errors can therefore move the model into states that are absent from offline recovery data. We introduce OnlineQAT, a two-stage framework that first obtains a usable low-bit initialization through block-wise QAT and then performs on-policy distillation (OPD) on student-generated responses. At each visited pre- fix, a frozen full-precision teacher provides a sampled reverse-KL training signal. On Qwen3-1.7B, Onl
קרא במקור המקורי