יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

T2SPO: אופטימיזציה של מדיניות ללמידת חיזוק

T2SPO: Trajectory-to-Step Policy Optimization for Agentic Reinforcement Learning
T2SPO היא שיטה לאופטימיזציה של מדיניות ללמידת חיזוק. היא משתמשת במסלולים קודמים כדי לספק משוב ללמידת מדיניות. T2SPO נבדקה עם מודלים בגודל 1.5B ו-7B והראתה שיפור בהצלחת משימות.
תקציר מקורי באנגליתarXiv:2610.00388v1 Announce Type: new Abstract: Reinforcement learning enables large language model (LLM) agents to learn multi-step behaviors through interaction with their environments. However, rewards in many interactive tasks reflect only the final outcome, providing limited guidance on which intermediate decisions advance the task. Successful training trajectories contain intermediate states that can provide supervision for subsequent interactions. We introduce Trajectory-to-Step Policy Optimization (T2SPO), a method that uses past interaction trajectories to provide step-level feedback for policy learning. T2SPO derives remaining-distance targets from successful trajectories and pairs them with representations of the states visited along the way. Conditioned on these examples, a pre
קרא במקור המקורי