יום שני, 5 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

פער כלליות ב-ProcGen

What Does a ProcGen Generalization Gap Measure? Action Rules, Residual Entropy, and the Missing Random Floor
חוקרים מציעים דרך למדוד פער כלליות ב-ProcGen, תוך שימוש במדד 'רצפת רנדומליות'. המחקר בוחן שמונה סביבות ProcGen עם אלגוריתם PPO. התוצאות מראות שהפער הוא משמעותי רק כאשר מולאם ברצפת רנדומליות.
תקציר מקורי באנגליתarXiv:2609.32532v2 Announce Type: replace Abstract: A generalization gap in reinforcement learning, return on training levels minus return on held-out levels, is usually reported without a reference point. We argue that it should be read against a measured random floor: the return of a uniform-random policy on the same levels under the same evaluation harness. On eight ProcGen environments with PPO at a compute-limited budget (8M steps, 16 parallel environments; three games extended to 25M), the floor changes what standard numbers mean. The test-time action rule decides which policy is measured: in miner, the sampled policy scores 5.1x the floor on held-out levels while its argmax scores below it in every run, and greedy evaluation places two environments significantly below the floor. Use
קרא במקור המקורי