כתבה
arXiv cs.LG ·
ARMOR: Stabilizing On-Policy LLM RL with Off-Policy Anchor Samples
תקציר מקורי באנגליתarXiv:2607.10481v2 Announce Type: replace Abstract: Reinforcement learning (RL) has significantly enhanced the reasoning capabilities of large language models (LLMs), yet the training process remains notoriously fragile. In this work, we investigate a critical source of this instability: over-optimization, where models exploit training heuristics at the expense of generalizable reasoning. While reverse KL regularization is the standard defense against such degradation, our analysis reveals that it is often insufficient in this regime, as it fails to ensure comprehensive coverage of the reference distribution. To address this, we propose ARMOR (Anchor Rollout and Mixed Optimization for RL), a framework that shifts the paradigm from passive penalty to active sample stabilization. ARMOR compr
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית