כתבה
arXiv cs.CL ·
למה גורמי LLM נכשלים בחקירת אזורים חדשים? עמדת תיאור עולם
Why Do LLM Agents Fail in Exploring New Environments? A World-Modeling Perspective
גורמי LLM נכשלים לשפר את עצמם באזורים חדשים עקב קריסת חקירה. חוקרים חקרו תיאור עולם כדי לפתור את הבעיה. התיאור נבנה כך שהוא יכול לשרת כהכנה ללמידה בעצמה. ניסויים מוצגים כדי להדגים את הצלחת התיאור.
תקציר מקורי באנגליתarXiv:2510.15047v2 Announce Type: replace-cross Abstract: Large Language Models (LLMs) as agents often fail to improve in new environments. We identify and characterize a failure mode we call exploration collapse: under reinforcement learning (RL) in environments whose states are unfamiliar to the policy, Pass@k, the probability that at least one of k sampled trajectories succeeds, drops markedly over training even as Pass@1 edges up, revealing increasingly brittle exploration; environments closer to the pretraining distribution show no such decline. We trace this collapse to weak grounding in environment states and dynamics, and study a simple remedy: explicitly teaching the agent to estimate the current state and predict its transitions before optimizing for reward. We instantiate it as
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית