יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

התאמת מודלים גנרטיביים חד-שלביים

Aligning One-Step Generative Models with Reward-Weighted Transport Distillation
RWTD היא שיטה חדשה להתאמת מודלים גנרטיביים חד-שלביים. השיטה משתמשת בתחבורה אופטימלית ורגרסיה קבועה. RWTD משפרת את הביצועים של מודלים גנרטיביים.
תקציר מקורי באנגליתarXiv:2609.30840v1 Announce Type: new Abstract: One-step generators enable high-quality visual generation with a single network evaluation, but their post-training is difficult: general implicit generators provide neither tractable likelihoods nor denoising trajectories, and many rewards are non-differentiable. We introduce Reward-Weighted Transport Distillation (RWTD), a post-training method that requires only generated samples and scalar reward evaluations. Rather than aligning solely to the conventional reward-tilted reference distribution, RWTD constructs an adaptive target that mixes separately tilted current and reference distributions. The current component incorporates improvements discovered during training, while the reference component anchors the target to the pretrained genera
קרא במקור המקורי