כתבה
arXiv cs.CL ·
Dynamic Optimizations of LLM Ensembles with Two-Stage Reinforcement Learning Agents
תקציר מקורי באנגליתarXiv:2502.04492v3 Announce Type: replace Abstract: The advancement of LLMs and their accessibility have triggered renewed interest in multi-agent reinforcement learning as robust and adaptive frameworks for dynamically changing environments. This paper introduces \texttt{RL-Focal}, a two-stage RL agent framework that routes and ensembles LLMs. \textit{First}, we develop the Decider RL-agent, which learns to dynamically select an ensemble of small size ($m_i$) among $N$ LLMs ($m_i \ll N$) for incoming queries from a user-defined downstream task $i$, by maximizing both error-diversity and reasoning-performance of the selected ensemble through iterative updates of task-adaptive rewards and policy. \textit{Second}, to enable effective fusion of dynamically selected LLMs, we develop the stage-
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית