יום חמישי, 8 באוקטובר 2026 LIVE
AI־INFO

כתבה arXiv cs.CL ·

LiveMACE: בדיקת יכולות של LLM בשוקים רב-תעשייתיים

LiveMACE: Process-Aware Evaluation of LLM Agent Capabilities in Evolving Markets
בדיקת יכולות של LLM בשוקים רב-תעשייתיים באמצעות שוקים פיננסיים חיים כמעבדה. המחקר ניצב על כנפיו של LiveMACEBench, תשתית של בדיקה שמשתמשת בשוקים פיננסיים חיים כמעבדה. המחקר כולל 5 LLM חדשים, כולל GPT-5, שהופעלו במשך 30 יום. המחקר מצא כי יש תפניות גדולות בין התוצאות לבין היכולות של ה-LLM. המחקר גם מצא כי יש תפניות גדולות בין הסגנון שבו ה-LLM פועל לבין התוצאות.
תקציר מקורי באנגליתarXiv:2610.09872v1 Announce Type: cross Abstract: Evaluating agents by outcomes alone can obscure the capabilities that produce them. This problem is especially pronounced in evolving environments, where outcomes reflect a closed-loop interaction between agent behavior and changing external conditions. We introduce LiveMACEBench, a process-aware benchmark that uses live financial markets as a naturally evolving testbed for persistent LLM agents. Five frontier LLMs operate along continuous trajectories under matched Tool Use, Persistent Memory, Rule Following, and Multi-Agent Collaboration configurations. We evaluate them through both realized outcomes and mechanism-specific diagnostics derived from complete decision traces. Across 30 days of live evaluation, we find a pronounced outcome-ca
קרא במקור המקורי