כתבה
arXiv cs.AI ·
מה נמצא בפרקסי: דירופאוט לטרקטוריות להטמעה על-פוליצי
What Shared Prefixes Hide: Trajectory Dropout for On-Policy Distillation
דירופאוט לטרקטוריות משפרת הטמעה על-פוליצי על ידי החלשת הסבילות של הטיפול באובייקטים.
תקציר מקורי באנגליתarXiv:2609.33455v2 Announce Type: replace Abstract: On-policy distillation (OPD) trains a student model on its own trajectories using dense token-level feedback from a stronger teacher model. Since each update is conditioned on the reasoning prefix already generated by the student, the prefix also shapes how effectively teacher feedback is converted into learning. We find that shared prefixes can lead to weak token-level updates, a phenomenon we call Prefix-Induced Supervision Attenuation (PISA). This attenuation arises in two common cases. (i) High student confidence can weaken corrective gradients even when the teacher disagrees. (ii) Tokens that rely on earlier reasoning can receive learning signals as weak as those for simple local continuations. To solve this problem, we propose Traje
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית