יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

אימון יועצים ל-LLM Agents מתוך תוצאות משימה

Training Advisors for LLM Agents from Task Outcomes
אימון יועצים ל-LLM Agents מתוך תוצאות משימה. Caddie, טכניקה לאימון ביקורת, משפרת את הצלחת Qwen3-4B ב-25% ומעברת את הביצועים של Kimi K3.
תקציר מקורי באנגליתarXiv:2610.09858v1 Announce Type: cross Abstract: Large language model agents tackle multi-step tasks by interleaving reasoning and tool calls with observations from the environment. Prior work has shown that natural-language feedback can help these agents revise their decisions during task execution. We introduce Caddie, a method for training critics to provide natural-language analysis and advice as agents work through a task. Unlike approaches that rely on step-level labels or reference critiques, Caddie learns from whether the agent ultimately succeeds after receiving the critic's feedback. We optimize the critic through reinforcement learning while keeping the base model frozen. Trained on multi-hop question answering with a single base model, our Qwen3-4B critic improves success rate
קרא במקור המקורי