כתבה
arXiv cs.AI ·
SAKI: ניהול סיוע מקסימלי-מקושר ללימוד תלמיד על-פוליצי
SAKI: Maximal-Coupling-Routed Teacher Supervision for On-Policy Distillation
SAKI מציגה טכניקה ללימוד תלמידים על-פוליצי, המשלבת רולאוט נחושה של המורה ומקסימלי-מקושר. הטכניקה משפרת את הביצועים של תלמידים ב-7 מבחני סיבוכיות.
תקציר מקורי באנגליתarXiv:2609.36601v1 Announce Type: new Abstract: On-policy distillation (OPD) reduces train-test state mismatch by training a student on its own generated trajectories, but weak students may visit teacher-misaligned prefixes where supervision is less representative. We introduce SAKI (Supervision Allocation with KL-constrained Interpolation), which combines a KL-constrained teacher-guided rollout with maximal coupling and reuses realized accept/correction events to route token-level supervision. Accepted positions retain sampled-token reverse-KL supervision, while correction positions receive direct supervision on the teacher's highest-probability token. Under maximal coupling, the correction probability is exactly TV(p_t, q_t), so the same trust-region radius controls rollout deviation and
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית