כתבה
arXiv cs.LG ·
LLMZero: גילוי של סטרטגיות אדפטיביות לאימון RL לאחר האימון דרך סוכני LLM
LLMZero: Discovering Adaptive Training Strategies for RL Post-Training via LLM Agents
LLMZero, סוכן אגנטיבי, גילה שסטרטגיות אימון אדפטיביות יכולות לשפר את הביצועים של RL עד 140%.
תקציר מקורי באנגליתarXiv:2606.18388v2 Announce Type: replace Abstract: RL post-training strategies are dataset-dependent and reveal a recurring empirical pattern: capacity parameters accumulate monotonically across stages, while regularization parameters predominantly oscillate in response to shifting training dynamics. This distinction highlights a potential flaw in fixed training schedules: by forcing all parameters along rigid paths, they fail to capture the dynamic exploration-exploitation tradeoffs that regularization must track. We uncover this through LLMZero, an agentic system that optimizes training trajectories via tree search by diagnosing pathologies at each checkpoint and proposing coordinated multi-parameter transitions. Across four diverse GRPO tasks, LLMZero discovers strategies that improve
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית