יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

למידת התנהגות עם פיצוי על מחלוקת לשם שליטה בקונטרול רצפי

Disagreement-Regularized Imitation Learning for Image-Based Continuous Control with Gaussian and Beta Policies
למידת התנהגות עם פיצוי על מחלוקת לשם שליטה בקונטרול רצפי. המחקר חקר את יעילות של DRIL בשליטה בקונטרול רצפי עם פיצוי על מחלוקת.
תקציר מקורי באנגליתarXiv:2609.38407v1 Announce Type: new Abstract: Purpose: Behavior cloning can accumulate errors when a learned controller visits states outside the demonstrated distribution. This study evaluates whether Disagreement-Regularized Imitation Learning (DRIL), which converts disagreement among cloned policies into a reinforcement-learning reward, improves image-based continuous control. Methods: A controlled CarRacing study combines Gaussian and Beta learner policies, demonstrations from either a clipped Gaussian expert or an intrinsically bounded Beta expert, one or 20 trajectories, deterministic and stochastic evaluation, and three retained stages: behavior cloning, the highest 10-episode training-score checkpoint, and the final DRIL checkpoint. The disagreement ensemble contains five Gaussia
קרא במקור המקורי