כתבה
arXiv cs.AI ·
Topology-Guided Modular Actor-Critic Learning for Continuous Systems under Temporal Objectives
תקציר מקורי באנגליתarXiv:2304.10041v4 Announce Type: replace Abstract: We study formal policy synthesis for continuous-state stochastic systems under linear temporal logic specifications. The product of the system with the automaton of the specification has a hybrid state space with sparse rewards. We introduce a generalized optimal backup order, defined in reverse to a topological order over automaton states, that guides value backups and provably preserves optimality. We further present a model-free actor-critic algorithm whose policy evaluation solves a constrained optimization problem by the augmented Lagrangian method, yielding hyperparameter self-tuning, and prove its optimality and convergence in the tabular case. Since integer encodings of automaton states impose a spurious ordinal relationship on fu
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית