כתבה
arXiv cs.LG ·
קידום רכיבי תכנון: לקחים מ-160,000 רצי טריינינג
JumpStart Your Policy Learning with Lessons from 160,000 Training Runs
מחקר גדול-היקף של למידת רכיבי תכנון ללא תקשורת, עם 160,000 רצי טריינינג. המחקר חושף כי רכיבי תכנון שונים יכולים להיות טובים בסביבות שונות, וכי תקינה של פרמטרי הרכיב יכולה לשנות את תוצאות הלמידה. המחקר גם מציע רכיב-מומלץ תנאי-סביבה, שמספק המלצות לרכיבי תכנון עבור פרקטיקנים.
תקציר מקורי באנגליתarXiv:2609.13730v1 Announce Type: new Abstract: Reliable progress in offline policy learning depends on careful reporting, well-tuned baselines, and evaluation across diverse conditions. Prior work has shown that results can be sensitive to reporting choices, hyperparameter tuning, and dataset properties, but these sources of variability have not been systematically investigated together at the scale needed to understand how they shape conclusions. To address this gap, we present a large-scale empirical study of offline reinforcement and imitation learning, training over 160,000 policies across 114 datasets. At this scale, no algorithm dominates: aggregate performance among the strongest methods is often close, but the leaders differ substantially across environments. We find that proper h
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית