כתבה
arXiv cs.LG ·
ReNFT: Repairing Mode Collapse in Reward Post-Training via Internal Probability-Mass Recalibration
תקציר מקורי באנגליתarXiv:2609.00061v2 Announce Type: replace Abstract: Reward post-training of diffusion generators inevitably concentrates probability mass on a few reward-favored modes, a mode collapse that erases within-prompt diversity. Existing methods for mitigating collapse rely on external signals or interfaces, augmenting the reward with perceptual objectives, adjusting reference regularization, or modifying the text encoder, but none repairs an adapter that has already collapsed while preserving the acquired reward. We observe that online post-training primarily reallocates probability mass over capabilities inherited from pretraining rather than learning new visual content. Collapse is therefore suppression, not deletion, and can be reversed from within the generator. We propose ReNFT, which repai
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית