כתבה
arXiv cs.CL ·
OPTS-TTPO: Enhancing Finite-Sample Policy-Gradient Learning with Tree Search
תקציר מקורי באנגליתarXiv:2609.40035v1 Announce Type: new Abstract: The policy-gradient theorem gives the exact gradient under the current policy, but finite on-policy samples may miss rare high-return trajectories. We study whether tree search improves their coverage within a fixed budget while controlling gradient bias. We introduce On-Policy Parallel Tree Search (OPTS) and Tree Trajectory Policy Optimization (TTPO) using on-policy tree trajectories, which sample new suffixes from the current policy at visited states. This needs no action-distribution correction, although branching changes state visitation. Our Branch Aggregation Lemma shows that branch-weighted tree statistics recover chain expectations when branch choices and weights are fixed before outgoing transitions are sampled. OPTS selects expansio
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית