כתבה
arXiv cs.AI ·
GIFT: Reconciling Post-Training Objectives via Variational Finite-Temperature Gibbs Initialization
תקציר מקורי באנגליתarXiv:2601.09233v3 Announce Type: replace-cross Abstract: The prevailing post-training paradigm for Large Reasoning Models (LRMs)---Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL)-suffers from an intrinsic optimization mismatch: the rigid likelihood maximization in SFT induces distributional collapse, thereby exhausting the exploration space necessary for subsequent RL. Motivated by the Gibbs optimum of KL-regularized RL, we derive a token-level variational surrogate that makes the SFT target structurally compatible with the subsequent RL stage, and propose Gibbs Initialization with Finite Temperature (GIFT). Standard SFT emerges as a degenerate zero-temperature limit of this surrogate, while a finite temperature preserves structural diversity. Our experiments demonstr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית