כתבה
arXiv cs.LG ·
StoreBench: A Live-Commerce Environment for Evaluating and Training Autonomous Operator Agents
תקציר מקורי באנגליתarXiv:2610.10942v1 Announce Type: cross Abstract: Reinforcement learning environments are now a primary lever for improving large language model (LLM) capabilities in post-training, yet most agentic benchmarks remain static: the world moves only when the agent acts, the reward is a terminal verdict, and the pass bar is set arbitrarily. We introduce StoreBench, a live-commerce environment in which an agent runs a mid-size online apparel store on a production-grade commerce backend, testing long-horizon planning and economic judgment under uncertainty. Customers order around the clock, suppliers reprice and fail, and market shocks arrive with partial or no warning. The agent acts through the same 29 merchant tools a human operator would use, under a windowed operation budget that makes simul
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית