כתבה
arXiv cs.AI ·
Best-of-$N$ Guidance for Test-time Diffusion Alignment
תקציר מקורי באנגליתarXiv:2610.05108v2 Announce Type: replace-cross Abstract: Diffusion models achieve strong generative performance but often struggle to align generated samples with human preferences measured by a reward model. A simple yet effective algorithm for test-time alignment is Best-of-$N$ (BoN) sampling, which draws $N$ i.i.d. samples from a pre-trained diffusion model and outputs the single highest-reward sample. Despite its empirical success, BoN makes limited use of reward information, as it is incorporated only at the final selection stage without influencing the reverse diffusion trajectory during sampling. Consequently, BoN sampling does not improve the average alignment of generated samples and is primarily suited to single-output settings. We propose Best-of-$N$ Guidance (BoNG), a novel me
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית