כתבה
arXiv cs.LG ·
ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement
תקציר מקורי באנגליתarXiv:2609.13425v1 Announce Type: new Abstract: Training diffusion models with multiple rewards requires distinguishing user preference from reward informativeness. User preference determines how much each reward should contribute to the overall objective; reward informativeness determines when its feedback is useful during denoising. Some rewards can meaningfully evaluate a sample as soon as global structure emerges, but others become informative only when the sample is nearly clean. To address both questions jointly, we propose ReCAST (Reward Credit ASsignment across T}imesteps), the first method, to our knowledge, for per-reward, timestep-dependent credit assignment in diffusion reward fine-tuning. ReCAST separates user preferences from temporal allocation through a reward-by-timestep w
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית