כתבה
arXiv cs.LG ·
PhantomEnvironments: Training LLM Agents in Fictional Worlds
תקציר מקורי באנגליתarXiv:2609.40221v1 Announce Type: new Abstract: Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environme
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית