יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

מדד רווח-קצב: פוליצי גרדיאנט להנדסת רשתות למידה

Reward-rate Policy Gradient for Efficient Machine Learning Engineering Agents
במאמר זה, המחברים מציגים פוליצי גרדיאנט חדש, המתפקד כמדד רווח-קצב, כדי לשפר את יעילות רשתות הלמידה. הם מציגים דוגמאות ובודקים את האפקטיביות של הפוליצי בשני סקנריות שונים.
תקציר מקורי באנגליתarXiv:2609.36393v1 Announce Type: new Abstract: Traditional reinforcement learning (RL) techniques focus on maximizing expected cumulative reward, where each action assumes to take a constant unit of time. However, this assumption does not hold for agentic RL tasks such as machine learning engineering (MLE) agents, where actions involve data loading, feature engineering, and model training that take variable durations. Efficiency matters in modern agentic RL where actions are costly. To address this limitation, we adapt from continuous-time RL and Semi-Markov Decision Process (SMDP) formulation and propose Reward-rate Policy Gradient (RPG), where we focus on optimizing the reward rate -- the long-term reward per unit of time. RPG estimates the reward rate from off-policy samples, then char
קרא במקור המקורי