כתבה
arXiv cs.LG ·
GIFT: התאמה של מטרות אחרי האימון דרך הגברת טמפרטורה סופית
GIFT: Reconciling Post-Training Objectives via Variational Finite-Temperature Gibbs Initialization
המאמר GIFT מציג פתרון לבעיית התאמה של מטרות אחרי האימון של מודלי LRM. הפתרון, GIFT, משתמש בהגברת טמפרטורה סופית כדי לשמר תכונות סטרוקטורליות. התוצאות המדעיות מציגות את יעילות GIFT בהשוואה לטכניקות אחרות.
תקציר מקורי באנגליתarXiv:2601.09233v3 Announce Type: replace Abstract: The prevailing post-training paradigm for Large Reasoning Models (LRMs)---Supervised Fine-Tuning (SFT) followed by Reinforcement Learning (RL)-suffers from an intrinsic optimization mismatch: the rigid likelihood maximization in SFT induces distributional collapse, thereby exhausting the exploration space necessary for subsequent RL. Motivated by the Gibbs optimum of KL-regularized RL, we derive a token-level variational surrogate that makes the SFT target structurally compatible with the subsequent RL stage, and propose Gibbs Initialization with Finite Temperature (GIFT). Standard SFT emerges as a degenerate zero-temperature limit of this surrogate, while a finite temperature preserves structural diversity. Our experiments demonstrate th
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית