כתבה
arXiv cs.CL ·
סביבות Phantom: אימון סוכנים LLM בעולמות בדיוניים
PhantomEnvironments: Training LLM Agents in Fictional Worlds
חוקרים פיתחו PhantomEnvironments, סביבות RL סינתטיות לאימון סוכנים LLM. הסביבות מאפשרות אימון סוכנים בעלויות נמוכות וללא צורך בנתונים מציאותיים. הסוכנים המאומנים מצליחים להעביר את הידע לסביבות אמיתיות.
תקציר מקורי באנגליתarXiv:2609.40221v1 Announce Type: cross Abstract: Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environ
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית