כתבה
arXiv cs.LG ·
Masked Diffusion Language Models are Strong and Steerable Text-Based World Models for Agentic RL
תקציר מקורי באנגליתarXiv:2607.16204v2 Announce Type: replace-cross Abstract: Recent growth in reinforcement learning (RL) has surfaced a need for diverse, specialized training environments. Hand-curated environments with fixed task and reward difficulties become ineffective signals as model performance improves, and sparse rewards over long horizons induce mode collapse on specific workflows or tool structures. World models that simulate environment states have matched pure rollout performance, making them promising for scaling diversity on-demand. However, autoregressive (AR) world models suffer from a left-to-right bias preventing conditioning on globally interdependent state anchors such as tool schemas, prior turns, and expected outcomes. We (i) formalize text-based world modeling as a steerable transiti
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית