כתבה
arXiv cs.LG ·
Double Descent and Malign Overfitting in Diffusion Models
תקציר מקורי באנגליתarXiv:2609.26392v2 Announce Type: replace Abstract: Conventional wisdom in deep learning holds that overparameterization---having more parameters $p$ than training samples $n$---is benign: larger models generalize better and, even without regularization, interpolating models generalize well, the test error following a double-descent curve. One might expect the same benign overfitting for diffusion models, whose training reduces to regression, i.e. to minimizing a quadratic score-matching loss. Yet the opposite is observed: overfitting here is catastrophic, driving the model into a memorization regime. We resolve this paradox by combining experiments on U-Nets trained on CelebA with a random-features model for which we derive closed-form learning curves. We show that with a fixed number $m$
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית