כתבה
arXiv cs.AI ·
ServeLearnBench: How Well Can Agents Self-Improve from Serving Experience?
תקציר מקורי באנגליתarXiv:2610.07792v1 Announce Type: cross Abstract: Large language model agents are increasingly deployed to perform complex tasks in real-world environments. However, the knowledge required for correct behavior in these environments is often implicit, undisclosed, and subject to change over time. Recent continual-learning harnesses seek to address this challenge by enabling agents to improve from serving experience. Yet the effectiveness and limitations of these methods are not yet well characterized. Existing benchmarks provide only partial coverage: some explicitly provide the target knowledge, others assume a static environment, and those that support continual adaptation remain limited in scale and knowledge diversity. To enable systematic evaluation, we formalize an evolving-environmen
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית