כתבה
arXiv cs.LG ·
Iterative Policy Refinement through Semantic Rollout Analysis
תקציר מקורי באנגליתarXiv:2610.01652v1 Announce Type: new Abstract: Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks show
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית