כתבה
arXiv cs.LG ·
Safe on Average, Unsafe in the Tail: When Is the Episodic-Cost Tail Controllable?
תקציר מקורי באנגליתarXiv:2610.09508v1 Announce Type: new Abstract: Safe reinforcement learning seeks policies that maximize return while satisfying constraints on cumulative cost. Most methods impose these constraints on expected episodic cost. Consequently, standard evaluations report mean episodic cost without characterizing how cost is distributed across episodes. A policy that satisfies the mean-cost criterion may therefore remain unsafe in its worst episodes. Mean-cost reporting neither identifies this tail violation nor shows whether it can be brought within budget while preserving return. In this work, we measure the episodic-cost tail using $\mathrm{CVaR}_{0.1}$, the average cost of the worst $10\%$ of episodes. We classify a policy as tail-safe when $\mathrm{CVaR}_{0.1}$ is within the safety budget.
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית