כתבה
arXiv cs.CL ·
Adaptive Mutual Distillation for Balanced Multi-Task Post-Training of Large Language Models
תקציר מקורי באנגליתarXiv:2610.02856v1 Announce Type: new Abstract: Multi-task post-training of large language models (LLMs) aims to improve performance across tasks with unequal amounts of training data. Existing methods focus primarily on balancing task contributions during single-model training. Different task-balancing strategies can produce models with complementary strengths, creating opportunities for mutual distillation. However, the usefulness of cross-model supervision can vary across tasks, transfer directions, and stages of training. We propose Adaptive Mutual Distillation (AMD), a collaborative post-training framework that jointly trains two models with different task-balancing strategies. AMD evaluates candidate adjustments to distillation weights through short training probes shared across task
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית