יום שישי, 31 ביולי 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

NormGuard: תיקון קיצורי משך תקין ב-Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning

NormGuard: Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning
NormGuard הוא פסילה של קיצורי משך תקין שמטרתה לשמור על תיקון הפרס המושג ב-Reward-Preserving Norm Constraints in Flow-Matching Reinforcement Learning. הפסילה נבדקה בשני בסיסי דגמים והיא הציגה תוצאות טובות בהשוואה לבסיסי דגמים שלא השתמשו בפסילה.
תקציר מקורי באנגליתarXiv:2606.27771v3 Announce Type: replace Abstract: Reinforcement learning (RL) post-training improves the reward alignment of flow-based generators, but often degrades perceptual quality in ways that are not captured by the reward proxy. We identify a simple structural signature of this drift: across three post-training methods (NFT, AWM, DPO), RL fine-tuning inflates the per-step velocity norm $\|v_\theta\|$ by $5\%$ to $15\%$ relative to the reference. A form of norm inflation has been studied in classifier-free guidance (CFG), where rescaling the velocity back to a reference norm at inference time can mitigate the resulting artifacts. However, this inference-time correction does not transfer cleanly to RL: rescaling $v_\theta$ to match $\|v_{\text{ref}}\|$ at inference time neither imp
קרא במקור המקורי