יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

CAST: טירונות עם יתרון סיבתי-מבנה-מאורגן לאימון דיפוזיה למודלי דיפוזיה

CAST: Causal Advantage-Structured Training with Spatially Grounded Compositional Rewards for Diffusion Models
CAST פותר את הבעיות של רכיבת השליטה האונליין למודלי דיפוזיה. המאמר עוסק בפיתוח CAST, שהוא שיטת טירונות עם יתרון סיבתי-מבנה-מאורגן לאימון מודלי דיפוזיה. CAST פותר את הבעיות של רכיבת השליטה האונליין למודלי דיפוזיה, כולל קביעת חלון ה-SDE, פיצוי רגישות, ואפקטיביות נמוכה של פיצוי. CAST נבחן על שני מודלי דיפוזיה חזקים, FLUX.2-dev ו-Qwen-Image-2512, ומציג תוצאות טובות יותר מאשר טירונות רגילה.
תקציר מקורי באנגליתarXiv:2609.39441v1 Announce Type: cross Abstract: Online reinforcement learning has been extended to flow matching for diffusion model (DM) image generation. However, this paradigm faces three limitations: (1) Window selection. Existing methods manually set the stochastic differential equation (SDE) sampling window, i.e., the denoising steps where exploration noise is injected. We instead determine it from each model's denoising trajectory. (2) Reward saturation. Current methods rely on scoring models trained on human annotations; we find that such scores are extremely high and nearly indistinguishable on the latest SOTA open-source DMs, making advantage estimation largely ineffective. (3) Sample inefficiency. A single scalar reward collapses different failure modes into almost identical s
קרא במקור המקורי