כתבה
arXiv cs.LG ·
On Language Drift during RLVR Post-Training
תקציר מקורי באנגליתarXiv:2610.02015v1 Announce Type: new Abstract: Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---and can potentially impair CoT monitorability---the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-tun
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית