כתבה
arXiv cs.AI ·
איזור אמון Q Adjoint Matching
Trust Region Q Adjoint Matching
TRQAM: תכנות איזור אמון ללמידת רפלקסיה רב-שלבית יציבה, כולל שימוש ב-Gemini ו-LangGraph
תקציר מקורי באנגליתarXiv:2605.27079v2 Announce Type: replace-cross Abstract: Off-policy reinforcement learning of pretrained flow policies remains challenging due to the instability of optimization arising from the multi-step sampling process. Recently, Q-learning with Adjoint Matching (QAM) addressed this by recasting policy improvement as a stochastic optimal control (SOC) problem guided by a learned critic. However, QAM inherits a fundamental fragility of critic-guided improvement, since small critic errors can be exponentially amplified and often lead to performance collapse. This paper introduces Trust Region Q Adjoint Matching (TRQAM), a stable off-policy fine-tuning algorithm that adaptively controls the path-space KL between the fine-tuned and pretrained policies through projected dual descent. Speci
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית