כתבה
arXiv cs.AI ·
האם סוכנים LLM מבצעים את התוכניות שהם מכריזים?
Do LLM Agents Execute the Plans They Declare? From Planning-Mode Declaration to Pattern-Specific Execution
חוקרים בדקו את הפער בין תכנון לביצוע של סוכנים LLM. הם מצאו שסוכנים כאלה לא תמיד מבצעים את התוכניות שהם מכריזים, ושיש צורך במנגנונים טובים יותר לבחירת תוכנית הטובה ביותר ולביצועה. החוקרים השתמשו במודלים כמו LLaMA ו- LangChain.
תקציר מקורי באנגליתarXiv:2609.38108v1 Announce Type: new Abstract: Large language models (LLMs) enable agents to solve long-horizon tasks by generating a plan and then executing it in an environment. However, successful planning requires two distinct capabilities: selecting an appropriate plan for the task and executing it faithfully. Existing planner--executor systems can fail at either stage, while final task success alone cannot distinguish selection from execution failures. We therefore study the Plan Declaration--Execution Gap and introduce Planning-as-Routing, where an LLM declares one of four planning modes: Predefined, Sequential, Hierarchical, or Search, and a deterministic router dispatches the task to the corresponding pattern-specific executor. Across four benchmarks and three LLMs, we find three
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית