כתבה
arXiv cs.LG ·
KDFlow: A User-Friendly and Efficient Knowledge Distillation Framework for Large Language Models
תקציר מקורי באנגליתarXiv:2603.01875v4 Announce Type: replace-cross Abstract: Knowledge distillation (KD) is widely used to compress and post-train large language models (LLMs), yet many existing frameworks execute teacher inference with the same training-oriented backend as student optimization, leading to suboptimal efficiency. In this paper, we propose KDFlow, a novel framework for LLM distillation that features a decoupled architecture and employs SGLang for teacher inference. KDFlow combines SGLang for teacher inference with PyTorch FSDP2 for student optimization, allowing each model to run on a backend tailored to its workload. To enable efficient full-vocabulary distillation in this decoupled architecture, KDFlow transfers the teacher's final hidden states via Ray's object store and recomputes teacher
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית