יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

משפחה מאוחדת של שערים לטוקנים לשילוב מדיניות

A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients
חוקרים מציגים משפחה מאוחדת של שערים לטוקנים לשילוב מדיניות, המאפשרת התאמה גמישה יותר של מודלים. המחקר בוחן את הישגי המשפחה הזו עם מודל Qwen3-32B כמורה ו-Qwen3-4B כתלמיד.
תקציר מקורי באנגליתarXiv:2609.11768v1 Announce Type: cross Abstract: Per-token gating of forward/reverse KL losses has become a standard technique for on-policy knowledge distillation (OPD), but existing methods such as EOPD (Jin et al., 2026) and ToDi (Jung et al., 2025) each fix a single gating signal and a single gating direction, and the two have never been compared directly. We introduce a four-coefficient parameterization lambda_t = sigma(a * h_t + b * u(x) + c + d * gap_t) in which direction-aligned proxies of EOPD and ToDi appear as one-dimensional (1D) restrictions, and which adds multi-channel composition and an explicit bias as further degrees of freedom. On TweetEval (Barbieri et al., 2020) emotion and hate, with a Qwen3-32B teacher and a Qwen3-4B student, configurations in the full family reach
קרא במקור המקורי