יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

MeanFlowAdvantage: שיפור תגמול יציב

MeanFlowAdvantage: Stable Reward Fine-Tuning for Few-Step Average-Velocity Generators
MeanFlowAdvantage הוא אלגוריתם חדש לשיפור תגמול יציב. הוא משתמש בפונקציית תגמול משוקללת ומאפשר יצירת דגמים טובים יותר. האלגוריתם נבדק על SD3.5-Medium והראה תוצאות טובות.
תקציר מקורי באנגליתarXiv:2609.37670v1 Announce Type: new Abstract: MeanFlow enables efficient few-step generation by predicting interval-average velocities, but this representation creates a mismatch for reward fine-tuning: existing advantage-based objectives are typically defined on instantaneous velocities or equivalent $x_0$-space predictions, whereas inference directly uses the learned average-velocity map. We introduce MeanFlowAdvantage, a signed advantage-weighted least-squares objective for average-velocity generators. Our key construction uses a shared, detached MeanFlow derivative correction to express the reward objective in prediction space while making rollout and reference regularization exact penalties on the average-velocity network deployed at inference. The resulting formulation preserves Me
קרא במקור המקורי