כתבה
arXiv cs.CL ·
Distilling What Matters: Confidence-Aware Selective Distillation for Large Language Models
תקציר מקורי באנגליתarXiv:2609.36734v1 Announce Type: new Abstract: Knowledge Distillation (KD) trains a smaller-capacity student model to imitate a larger-capacity teacher model by matching output distributions, implicitly assuming the teacher to be a reliable oracle. In large language models (LLMs), this assumption often fails: teacher predictions can exhibit high entropy and hallucinations, causing standard KD to degrade well-calibrated student priors. We propose CaRE-KD, a confidence-gated distillation framework that replaces static objectives with uncertainty-adaptive optimization. CaRE-KD has two components: a token-level loss (CaRE-Divergence) that adaptively switches between Forward and Reverse KL divergence based on teacher--student confidence, and a batch-level epistemic rejection mechanism (Revival
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית