כתבה
arXiv cs.AI ·
PreviewDiff: Multimodal Critic-Guided Search over Diffusion Latents
תקציר מקורי באנגליתarXiv:2609.36199v1 Announce Type: cross Abstract: Diffusion models can produce striking images and videos, but they still struggle with the compositional details that make a generation faithful to a prompt, such as object counts, attribute binding, spatial relations, and temporally grounded actions. A common way to improve prompt satisfaction is to spend more compute at test time through Best-of-N sampling, but final-sample selection is fixed. Best-of-N can only choose among completed outputs and cannot repair a promising trajectory before it fails. We introduce PreviewDiff, a training-free test-time search method that turns diffusion sampling from scalar search into a multimodal critic-guided search over intermediate latents. At selected denoising checkpoints, PreviewDiff decodes a partia
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית