יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

PhGPO: אופטימיזציה מונחה פרומונים לתכנון כלי לאורך זמן

PhGPO: Pheromone-Guided Policy Optimization for Long-Horizon Tool Planning
במאמר זה, נוסחנאות PhGPO הוצגו, שמשתמשות בפרומונים כדי לאופטימיזציה של תכנון כלי לאורך זמן. השיטה הזו יכולה לשפר את יעילות התכנון של כלי ולהפחית את הזמן הדרוש לתכנון.
תקציר מקורי באנגליתarXiv:2602.13691v2 Announce Type: replace Abstract: Recent advancements in Large Language Model (LLM) agents have demonstrated strong capabilities in executing complex tasks through tool use. However, long-horizon multi-step tool planning is challenging, because the exploration space suffers from a combinatorial explosion. In this scenario, even when a correct tool-use path is found, it is usually considered an immediate reward for current training, which would not provide any reusable information for subsequent training. In this paper, we argue that historically successful trajectories contain reusable tool-transition patterns, which can be leveraged throughout the whole training process. Inspired by ant colony optimization where historically successful paths can be reflected by the phero
קרא במקור המקורי