יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

משפחה מאוחדת של שערים לטוקנים לעיבוד ידע

A Unified Per-Token Gating Family for On-Policy Distillation: FKL/RKL Mixing with Multi-Channel and Bias Coefficients
חוקרים מציגים משפחה מאוחדת של שערים לטוקנים לעיבוד ידע, המאפשרת התאמה מותאמת יותר של מודלים. המחקר בוחן את הישגי המשפחה עם מודל Qwen3-32B כמורה ו-Qwen3-4B כתלמיד.
תקציר מקורי באנגליתarXiv:2609.11768v1 Announce Type: new Abstract: Per-token gating of forward/reverse KL losses has become a standard technique for on-policy knowledge distillation (OPD), but existing methods such as EOPD (Jin et al., 2026) and ToDi (Jung et al., 2025) each fix a single gating signal and a single gating direction, and the two have never been compared directly. We introduce a four-coefficient parameterization lambda_t = sigma(a * h_t + b * u(x) + c + d * gap_t) in which direction-aligned proxies of EOPD and ToDi appear as one-dimensional (1D) restrictions, and which adds multi-channel composition and an explicit bias as further degrees of freedom. On TweetEval (Barbieri et al., 2020) emotion and hate, with a Qwen3-32B teacher and a Qwen3-4B student, configurations in the full family reach hi
קרא במקור המקורי