יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

גנרליזציה חלשה-חזקה עם דיסטילציה הפוכה

Eliciting Weak-to-Strong Generalization with On-Policy Reverse Distillation
On-Policy Reverse Distillation (OPRD) היא שיטה חדשה לגנרליזציה חלשה-חזקה. OPRD מאפשרת למודלים חזקים ללמוד ממודלים חלשים ולעקוף אותם. השיטה משתמשת בדיסטילציה הפוכה כדי לחזק את הלמידה של המודל החזק.
תקציר מקורי באנגליתarXiv:2609.08798v1 Announce Type: new Abstract: Weak-to-strong generalization asks whether stronger models can learn from weaker supervisors and surpass them. This question is particularly important for successive model generations and multi-domain consolidation, where repeating frontier-scale post-training from scratch can be prohibitively expensive. Yet conventional distillation treats the weak teacher as an optimization target, potentially imposing its capacity ceiling on the student. We introduce On-Policy Reverse Distillation (OPRD), which evaluates the teacher's policy shift relative to its reference policy on student rollouts and amplifies the component of the student's verifier-driven policy gradient along that direction. By rescaling only verifier-supported updates, OPRD preserves
קרא במקור המקורי