יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

LeanGRPO: פינוי חישוב חוזר ב-Diffusion RL

LeanGRPO: Eliminating Redundant Recomputation in Diffusion RL
LeanGRPO מפנה חישוב חוזר ב-Diffusion RL. המאמר מציג שני תוכניות אימון חדשות שמאפשרות חישוב חוזר חסכוני.
תקציר מקורי באנגליתarXiv:2609.03528v1 Announce Type: new Abstract: Diffusion reinforcement learning (RL) has recently achieved significant success in post-training image and video generative models. However, most diffusion RL methods, including DanceGRPO and FlowGRPO, recompute selected timesteps with gradient tracking after rollout. Under on-policy training with the same backend for rollout and update, this recomputation is mathematically redundant. Intuitively, the rollout and policy update steps can reuse the same feed-forward backbone to avoid redundant computation, but doing so can incur a large memory overhead during rollout. To address the issue, we present LeanGRPO by restructuring the data-parallel layout and introducing two recompute-free training schedules for trajectory-logprob diffusion RL: (1)
קרא במקור המקורי