יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

JAGG: Jacobian-Aggregated Group Gradient לאופטימיזציה יעילה של GRPO למודלי דיפוזיה

JAGG: Jacobian-Aggregated Group Gradient for Efficient GRPO Training of Diffusion Models
JAGG מציע אופטימיזציה יעילה של GRPO למודלי דיפוזיה, עם עיכוב של 2x בפענות האחורה. החידוש, JAGG, משתמש באינטרפולציה של יעקביאנים כדי לאגגרגטר גרדיאנטים. התוצאות המדעיות מציגות עיכוב של 2x בפענות האחורה, עם ירידה זניחה באיכות. הקוד לעבודה זו ניתן לגישה דרך https://github.com/SchumiDing/JAGG.
תקציר מקורי באנגליתarXiv:2607.17572v3 Announce Type: replace Abstract: Group Relative Policy Optimization (GRPO) is a powerful reinforcement learning algorithm for aligning generative models with human preferences. While successful in large language models~\cite{shao2024deepseekmathpushinglimitsmathematical}, its extension to diffusion and flow matching models introduces a severe computational bottleneck: gradients must be back-propagated through the high-capacity DiT backbone at \emph{every} timestep of the sampling trajectory, making high-resolution text-to-image (T2I) training prohibitively expensive. Training-free DiT inference acceleration methods (e.g., $\Delta$-DiT, ScalingCache) exploit the fact that DiT hidden states and velocity predictions vary \emph{smoothly and nearly linearly} along the traject
קרא במקור המקורי