כתבה
arXiv cs.AI ·
הטבעת של רכיבת הלמידה: פוסט-אימון של LLM
Distilled Reinforcement Learning for LLM Post-training
מחקר חדש: Distilled Reinforcement Learning (DRL) משפר את התפיסה והתאמה של LLM לאחר האימון. המחקר מציע שיטה חדשה להעברת ידע ממודלי LLM אחד לאחר. השיטה, שנקראת DRL, משלבת שיטות של רכיבת הלמידה והעברת ידע. המחקר נערך על ידי צוות מדעניות ומדעני טכנולוגיה מאוניברסיטת [שם האוניברסיטה].
תקציר מקורי באנגליתarXiv:2607.17247v2 Announce Type: replace-cross Abstract: Large language model (LLM) post-training is essential for improving reasoning, adaptation, and alignment. Existing methods mainly follow two paradigms: reinforcement learning (RL) and on-policy distillation (OPD). However, RL relies on coarse-grained outcome supervision, resulting in difficult credit assignment and limited capability to acquire new knowledge. OPD, meanwhile, unconditionally matches teacher logits through KL divergence, which creates a dilemma: similar teachers provide little new knowledge, while substantially different teachers often yield ineffective guidance, largely restricting OPD to within-family distillation. We propose Distilled Reinforcement Learning (Distilled RL), which integrates teacher supervision into
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית