יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

שליטה נגדית קבועה להתפשרות שחור-קופסה

Persistent Negatives for Adversarial Black-Box On-Policy Distillation
שליטה נגדית עם שליטה קבועה נגדית משפרת התפשרות שחור-קופסה באופן-מדידה.
תקציר מקורי באנגליתarXiv:2609.30864v1 Announce Type: new Abstract: Black-box On-Policy Distillation (OPD) seeks to improve a student from its own generations when the teacher provides sampled responses but not token probabilities. Adversarial distillation offers one route: it learns a discriminator over prompt-matched teacher and student responses and uses its score as the policy reward. However, sampling discriminator negatives from the latest student at each step couples the learned reward to a negative distribution that changes after every policy update. We address this moving-target problem with persistent-negative adversarial distillation, a live-pool method that replaces a fraction of each discriminator batch with historical, prompt-matched teacher--student comparisons. Under matched discriminator comp
קרא במקור המקורי