כתבה
arXiv cs.AI ·
למידת Q-ערך חסרת סיכונים: דוח תצפית קצר-סטפ
Mini-Batch Risk-Averse Deep Q-Learning: A Robot Navigation Case Study
למידת Q-ערך חסרת סיכונים לניווט רובוטי. המאמר עוסק בפיתוח שיטה ללמידת Q-ערך חסרת סיכונים לניווט רובוטי. השיטה משתמשת בטכניקה של Q-למידה חסרת סיכונים ומתמקדת בביצוע ניווט רובוטי חסר סיכונים. המאמר כולל תיאור של השיטה, תיאור של המודלים והכלים שהושקעו בפיתוחה, ותיאור של התוצאות שהושגו באמצעותה.
תקציר מקורי באנגליתarXiv:2609.07998v1 Announce Type: new Abstract: We study the control of Markov decision processes in which the quality of a policy is evaluated by a dynamic, time-consistent Markov risk measure rather than by an expected discounted cost. The main obstacle to combining such measures with reinforcement learning is that a transition risk mapping depends on the transition kernel in a nonlinear way, and therefore cannot be estimated from a single observed transition. We remove this obstacle by employing mini-batch transition risk mappings: the mapping is applied to the empirical measure of $N$ independent next-state samples, and the result is averaged. The resulting mapping is again coherent. However, as an expected value of a function of $N$ next-state values, it admits an unbiased one-sample
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית