כתבה
arXiv cs.AI ·
על דחיפה של שפה במהלך RLVR לאחר הכשרה
On Language Drift during RLVR Post-Training
במאמר זה נחקרה תופעת דחיפה של שפה במהלך RLVR לאחר הכשרה. נראה שהתופעה קשורה לאופטימיזציה של RLVR ולא להכשרה סופרוויזד.
תקציר מקורי באנגליתarXiv:2610.02015v1 Announce Type: cross Abstract: Recent advances in LLM reasoning models---driven primarily by the paradigm of post-training via reinforcement learning with verifiable reward (RLVR)---have enabled them to accomplish impressively complex tasks. However, in parallel with their rising capabilities, LLMs have increasingly displayed signs of language drift in their chains of thought (CoTs): unusual, non-standard, and seemingly nonsensical language use. Although it is well-documented---and can potentially impair CoT monitorability---the causes of language drift are thus far poorly understood. In this paper, we identify the conditions under which language drift occurs: we prove theoretically that RLVR optimization pressure permits unbounded language drift, while supervised fine-t
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית