יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

שיקוף יכולות כלליות של LLM עם שימור תחום

Counteraction-Aware Multi-Teacher On-Policy Distillation for General Capability Recovery with Domain Preservation
שיקוף יכולות כלליות של LLM עם שימור תחום. המחקר מציע שיטה לשיקוף יכולות כלליות של LLM עם שימור תחום. השיטה, שנקראת CaMOPD, משתמשת בשיטת On-Policy Distillation כדי לשקף יכולות כלליות של LLM עם שימור תחום. CaMOPD משתמשת בשיטת Alternating Training כדי לשקף יכולות כלליות של LLM עם שימור תחום. CaMOPD משתמשת בשיטת Gap-Based Sample Selection כדי לשקף יכולות כלליות של LLM עם שימור תחום.
תקציר מקורי באנגליתarXiv:2605.27115v2 Announce Type: replace Abstract: Domain specialization can improve LLM behavior, but often weakens the general capabilities inherited from the original model. Recent Multi-Teacher On-Policy Distillation (MOPD) pipelines recover model capabilities by supervising student-generated trajectories with teacher feedback, but typically assume teacher-aligned prompt coverage, requiring prompts to match the teachers' training distributions. This assumption is difficult to satisfy when the general teacher is an open-source model whose post-training data are unknown. Instead of attempting to reconstruct this hidden distribution, we study general capability recovery with readily available proxy general prompts. We identify two failure modes of vanilla MOPD in this incomplete-coverage
קרא במקור המקורי