כתבה
arXiv cs.LG ·
Conditional Residual Prediction: Improving Autoregressive Video Diffusion without a Bidirectional Teacher
תקציר מקורי באנגליתarXiv:2610.11479v1 Announce Type: cross Abstract: Causal video diffusion models generate video autoregressively, which suits streaming, interactive, and long-video generation. Under standard training, however, they often yield lower generation quality than bidirectional models of the same size. Many existing approaches address this gap by initializing from or distilling a pretrained bidirectional teacher. We instead train a causal model from an image-model initialization, with no bidirectional video model at any stage. Because this path requires neither a large bidirectional teacher nor a complex distillation pipeline, it is simpler and more scalable. On this path, we find that a causal model trained on ground-truth history becomes strongly dependent on it, so that at inference errors in i
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית