כתבה
arXiv cs.LG ·
The Distillation Game: Adaptive Evaluations & Efficient Defenses
תקציר מקורי באנגליתarXiv:2605.22737v4 Announce Type: replace Abstract: Distillation attacks create a deployment trade-off for model providers: the same outputs that make a model more useful can also make it easier to imitate. We study this trade-off through a minimax game between a utility-constrained teacher and an adaptive student. Our framework yields tractable one-sided response rules: an adaptive evaluation rule in which the student reweights high-value examples, and a teacher-side defense template that suppresses outputs most useful for distillation. From a cheap proxy for example value, we derive Product-of-Experts (PoE), a simple forward-pass-only defense that combines the teacher with a proxy student during generation. Empirically, adaptive evaluation reveals a large passive--adaptive gap: on state-
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית