כתבה
arXiv cs.AI ·
STRAT: משימה עזרית בלמידת חיזוק
State Trace Rationale As Auxiliary Task in Reinforcement Learning
STRAT היא משימה עזרית שמאמנת סוכנים של למידת חיזוק לחזות תיאור קצר של מצבם. השיטה מוסיפה ראש עזרי אחד למדיניות סטנדרטית. STRAT מצליחה לפתור סביבות מורכבות שבהן למידת חיזוק סטנדרטית נכשלת.
תקציר מקורי באנגליתarXiv:2609.36867v1 Announce Type: new Abstract: We propose STRAT, an auxiliary task that trains deep reinforcement learning (RL) agents to predict a short textual trace of their own state. Inspired by human spatial navigation, the description combines landmark, route, and survey knowledge, tracking the agent's position, inventory, goals, and immediate progress. Environment rules generate this text online without human labelling. Our method adds a single auxiliary head to a standard policy. Across 60 sparse-reward XLand-MiniGrid tasks, STRAT solves complex environments where standard RL fails outright, while compacting state representations and preventing rank collapse. Beyond performance gains, the predicted trace provides a readable account of agent beliefs at every step for no extra cost
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית