כתבה
arXiv cs.AI ·
ReCAST: שיטה חדשה להקצאת קרדיט בדיפוזיה
ReCAST: Reward Credit Assignment across Timesteps for Online Diffusion Reinforcement
ReCAST היא שיטה חדשה להקצאת קרדיט בדיפוזיה, המאפשרת להבדיל ב间 user preference מרווח המידע של הפרס. השיטה משתמשת במטריצת משקלות כדי להקצות קרדיט לכל פרס בכל צעד דיפוזיה.
תקציר מקורי באנגליתarXiv:2609.13425v1 Announce Type: cross Abstract: Training diffusion models with multiple rewards requires distinguishing user preference from reward informativeness. User preference determines how much each reward should contribute to the overall objective; reward informativeness determines when its feedback is useful during denoising. Some rewards can meaningfully evaluate a sample as soon as global structure emerges, but others become informative only when the sample is nearly clean. To address both questions jointly, we propose ReCAST (Reward Credit ASsignment across T}imesteps), the first method, to our knowledge, for per-reward, timestep-dependent credit assignment in diffusion reward fine-tuning. ReCAST separates user preferences from temporal allocation through a reward-by-timestep
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית