כתבה
arXiv cs.LG ·
Learning What to Distill: Bilevel Top-K Token Selection for Self-Distillation in Large Language Models
תקציר מקורי באנגליתarXiv:2610.07247v1 Announce Type: new Abstract: Large language models have shown strong reasoning capabilities, but their high inference costs make knowledge distillation an important approach for transferring such capabilities to compact models in resource-constrained scenarios. On-policy self-distillation further reduces the reliance on external large teacher models while improving the reasoning ability of compact language models. However, existing methods typically either distill all token positions uniformly or select tokens using fixed heuristic criteria, assigning the same distillation strength to the selected positions rather than adaptively learning which tokens are most beneficial for distillation. To address these limitations, we propose BiToK-SD (Bilevel Top-K Token Selection fo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית