כתבה
arXiv cs.AI ·
Post-Training at the Edge of Detectability: A Game-Theoretic Approach to Fine-Tuning
תקציר מקורי באנגליתarXiv:2607.26358v1 Announce Type: cross Abstract: Reinforcement learning (RL) fine-tuning is widely used in language model training to improve model performance on a target task while limiting drift from a reference policy. A standard way to balance this trade-off is via a KL-regularized RL objective, although this formulation does not by itself provide a principled way to set the regularization coefficient. In practice, the coefficient is typically chosen heuristically or via hyperparameter search, which can lead to unnecessary overhead in training cost or undesirable reward-retention trade-offs. We instead propose a game-theoretic framework that gives this trade-off an explicit statistical interpretation. Specifically, we study a sequential game in which an agent chooses a policy to maxi
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית