יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

QATFactory: פלטפורמה ניידת לאימון-מודלים-לשוני-גדולים עם תכנון-מחדש-למידע-קצר

QATFactory: A Versatile, Deployment-Aligned Framework for Quantization-aware Training and Distillation of LLMs
פלטפורמה חדשה לאימון-מודלים-לשוני-גדולים עם תכנון-מחדש-למידע-קצר, כולל שימוש ב-QATFactory וב-Qwen. הפלטפורמה מאפשרת אימון-מודלים-לשוני-גדולים עם תכנון-מחדש-למידע-קצר, ומשפרת את תפעול-המודל-במצב-הפעלה.
תקציר מקורי באנגליתarXiv:2609.39223v1 Announce Type: new Abstract: Large language model (LLM) inference is increasingly moving toward lower precision to realize the throughput of hardware accelerators, but aggressive post-training quantization (PTQ) can degrade model quality. We present QATFactory, an open-source framework for deployment-aligned quantization-aware distillation (QAD) and reinforcement learning (QARL). QATFactory simulates deployment-time quantization while performing matrix multiplications in BF16, allowing models to adapt to quantization noise without requiring training hardware that natively supports the target format; for example, it supports NVFP4 training on H100 GPUs, which lack FP4 Tensor Cores. The framework supports NVFP4, MXFP4, and llama.cpp's Q4_K format; dense and mixture-of-expe
קרא במקור המקורי