כתבה
arXiv cs.AI ·
ARCO: Adaptive Rubrics with Co-Evolution for Multi-Step LLM-Based Agents
תקציר מקורי באנגליתarXiv:2606.21262v2 Announce Type: replace Abstract: Reinforcement learning for multi-step LLM agents often relies on scalar rewards that indicate success but cannot explain why a trajectory is good or bad. Rubric-based rewards improve interpretability through natural-language criteria, but existing methods share two limitations: they score at the trajectory level, offering no guidance for individual steps; and their scorer is closed-source and static, so it cannot adapt as the agent evolves during training. We propose ARCO (Adaptive Rubric CO-evolution), which generates a per-step rubric and predicts a rubric-conditioned step-level reward for each action, and continually updates this rubric model on on-policy rollouts so that its criteria and scores co-evolve with the agent's improving beh
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית