כתבה
arXiv cs.LG ·
גרדיאנטים של מדידיות-מרבית: פוליצי-גרדיאנטים שמצביעים על רכישת נתונים-מועטים
Multi-Fidelity Policy Gradients Stabilize Data-Scarce Reinforcement Learning
פוליצי-גרדיאנטים של מדידיות-מרבית: פתרון ללמידה-עצמית על-פי-נתונים-מועטים. ניתן להשיג תוצאות טובות יותר בלמידה-עצמית על-פי-נתונים-מועטים, כאשר נשתמש בפוליצי-גרדיאנטים של מדידיות-מרבית. פוליצי-גרדיאנטים של מדידיות-מרבית נותנים תוצאות טובות יותר בלמידה-עצמית על-פי-נתונים-מועטים, כאשר נשתמש בפוליצי-גרדיאנטים של מדידיות-מרבית.
תקציר מקורי באנגליתarXiv:2610.02505v1 Announce Type: new Abstract: Policy gradient methods for on-policy reinforcement learning (RL) can become unstable when expensive, scarce target-domain data yield noisy gradient estimates. We address this challenge by complementing limited high-fidelity (HF) target-domain data with abundant, cheap, but biased low-fidelity (LF) data, e.g., from a simplified simulator. Most existing methods directly optimize biased objectives based on LF data. In contrast, the recently introduced multi-fidelity policy gradient (MFPG) framework uses LF data solely to construct a control variate that reduces variance and improves HF data efficiency without biasing the policy gradient estimator. However, published work on MFPG is limited to REINFORCE on small-scale simulation tasks. We develo
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית