כתבה
arXiv cs.AI ·
PhantomEnvironments: טריינינג LLM Agents בעולמות רצויים
PhantomEnvironments: Training LLM Agents in Fictional Worlds
במאמר זה, נראה כיצד ניתן לטריינינג LLM Agents בעולמות רצויים, תוך שימוש באינטראקציות סינתטיות. התוצאות היו תוצאות טובות, וה-LLM Agents המטורנים הצליחו להעביר עצמם לבנצ'מארקים ריאל-עולם.
תקציר מקורי באנגליתarXiv:2609.40221v1 Announce Type: cross Abstract: Training LLM agents with reinforcement learning (RL) is bottlenecked by environments, which must provide verifiable rewards, support long-horizon interaction, and scale cheaply. Existing approaches rely on costly human-curated data or on LLM-generated environments that risk hallucinations and benchmark contamination. We show that LLMs can instead be trained into capable search agents using synthetic environments generated entirely by rules, whose generation requires no LLM and has zero marginal cost. We build PhantomEnvironments, multi-turn RL environments from fictional worlds, where agents must search a corpus of templated articles to answer multi-hop questions. Despite sharing no facts with the real world, these strikingly simple environ
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית