יום שלישי, 15 בספטמבר 2026 LIVE
AI־INFO

כתבה arXiv cs.LG ·

AgentRM: השיפור של כלליות של סוכנים עם תיאור פרסום

AgentRM: Enhancing Agent Generalization with Reward Modeling
AgentRM הוא פרויקט שמטרתו לשפר את כלליות של סוכנים עם תיאור פרסום. הפרויקט כולל שלושה גישות שונות לבניית תיאור פרסום, כולל תיאור פרסום מובהק, תיאור פרסום בלתי מובהק ו-LLM-as-a-judge.
תקציר מקורי באנגליתarXiv:2502.18407v2 Announce Type: replace-cross Abstract: Existing LLM-based agents have achieved strong performance on held-in tasks, but their generalizability to unseen tasks remains poor. Hence, some recent work focus on fine-tuning the policy model with more diverse tasks to improve the generalizability. In this work, we find that finetuning a reward model to guide the policy model is more robust than directly finetuning the policy model. Based on this finding, we propose AgentRM, a generalizable reward model, to guide the policy model for effective test-time search. We comprehensively investigate three approaches to construct the reward model, including explicit reward modeling, implicit reward modeling and LLM-as-a-judge. We then use AgentRM to guide the answer generation with Best-
קרא במקור המקורי