כתבה
arXiv cs.CL ·
AgentRM: שיפור כלליות הסוכנים עם מודל רווח
AgentRM: Enhancing Agent Generalization with Reward Modeling
AgentRM הוא מודל רווח כללי שמשפר את כלליות הסוכנים. הוא משמש כמדריך למודל מדיניות ומשפר את התוצאות במגוון משימות. המודל נבדק על מודל LLaMA-3-70B והראה שיפור משמעותי.
תקציר מקורי באנגליתarXiv:2502.18407v2 Announce Type: replace Abstract: Existing LLM-based agents have achieved strong performance on held-in tasks, but their generalizability to unseen tasks remains poor. Hence, some recent work focus on fine-tuning the policy model with more diverse tasks to improve the generalizability. In this work, we find that finetuning a reward model to guide the policy model is more robust than directly finetuning the policy model. Based on this finding, we propose AgentRM, a generalizable reward model, to guide the policy model for effective test-time search. We comprehensively investigate three approaches to construct the reward model, including explicit reward modeling, implicit reward modeling and LLM-as-a-judge. We then use AgentRM to guide the answer generation with Best-of-N s
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית