יום ראשון, 4 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.AI ·

שיפור מדיניות באמצעות ניתוח סמנטי

Iterative Policy Refinement through Semantic Rollout Analysis
חוקרים הציגו שיטה חדשה לשיפור מדיניות בלמידת חיקוי. השיטה משתמשת בניתוח סמנטי של ריצות מדיניות כדי לזהות תת-אופטימליות ולתקנן. השיטה הראתה שיפור של עד 15% בביצועים לעומת שיטות קודמות.
תקציר מקורי באנגליתarXiv:2610.01652v1 Announce Type: cross Abstract: Structured policies improve efficiency, robustness, and interpretability in imitation learning by introducing task-specific inductive bias, but existing structure generation methods rely either on extensive human input or on static domain knowledge encoded in LLMs, which may be inconsistent with the expert demonstrations. We propose a closed-loop framework that iteratively refines structured policies using LLM-guided analysis of policy rollouts. By logging rollouts as semantically meaningful tabular data and prompting the LLM to generate diagnostic analysis code, our method identifies suboptimalities in the policy structure and iteratively corrects them without requiring human instruction. Experiments on car racing and door opening tasks sh
קרא במקור המקורי