כתבה
arXiv cs.LG ·
ReTaCo: שליטה על-מטרה עם ניצור-תגובה תלויה במטרה
ReTaCo: Residual-Target Control for On-Policy Distillation
ReTaCo: שליטה על-מטרה עם ניצור-תגובה תלויה במטרה. המאמר עוסק בשיפור של On-Policy Distillation (OPD) על ידי ניצור-תגובה תלויה במטרה. ReTaCo (Residual-Target Control) הוא פרוטוקול חדש שמשלב ניצור-תגובה תלויה במטרה עם ניצור-תגובה תלויה במטרה. הפרוטוקול נועד לשפר את היכולת של OPD ללמד סטודנטים להפיק תגובות טובות יותר.
תקציר מקורי באנגליתarXiv:2609.39275v1 Announce Type: new Abstract: On-policy distillation (OPD) trains a student on its own generated prefixes with token-level teacher feedback, but transmitting or storing the teacher's full-vocabulary distribution at every token is costly. Entropy-aware OPD (EOPD) adds forward supervision to reverse KL to help the student recover plausible tokens it underestimates, using only the teacher's top-$k$ probabilities to limit cost. Because EOPD renormalizes these probabilities, its target assigns no mass to the omitted vocabulary. We prove that the resulting loss keeps pushing the student's top-$k$ mass toward one even after the student matches the teacher's relative probabilities within the top-$k$ set, so the teacher itself is not a stationary point whenever the omitted tokens
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית