כתבה
arXiv cs.AI ·
EnvACE: Internalizing Environment Dynamics via World Rehearsal for Agentic Reinforcement Learning
תקציר מקורי באנגליתarXiv:2608.06197v2 Announce Type: replace Abstract: Training large language model agents for long-horizon tool use typically relies on interactions with real or synthesized executable environments, whose construction and verification are costly, or on external simulators that are difficult to ground. We introduce EnvACE, an agentic reinforcement learning method that replaces external environment interaction during training with world rehearsal. The policy alternates between acting and rehearsal: it first generates a tool call, then plays the role of the environment to produce the response induced by that action, and conditions subsequent decisions on the rehearsed response. Both roles are jointly optimized end-to-end using task-success rewards. Through world rehearsal, the policy internali
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית