כתבה
arXiv cs.LG ·
On-Policy Parameter Update Direction Underlies Generalization in LLM Post-Training
תקציר מקורי באנגליתarXiv:2609.36659v1 Announce Type: new Abstract: The strong generalization performance of on-policy post-training paradigms has motivated studies of their parameter update behaviors. However, these studies treat the observed behaviors only as byproducts in on-policy training, overlooking their potential to serve as optimization principles for improving the generalization of other paradigms such as supervised fine-tuning (SFT). To address this limitation, we investigate whether there exists a specific on-policy update behavior that can achieve such improvements. First, our analyses reveal that SFT updates parameters along consistent directions, while the on-policy paradigm continuously adjusts the direction during training. This difference inspires us to focus on the cumulative update direct
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית