כתבה
arXiv cs.LG ·
MintEval: בדיקת אם מודלים גדולים מיישמים אסטרטגיות סחר
MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code
MintEval הוא בנך' לבדיקת יכולתם של מודלים גדולים ליישם אסטרטגיות סחר. הבנך' בודק אם הקוד שנכתב על ידי המודל מיישם את הלוגיקה שתוארה. Claude Opus 5.5 השיג תוצאות טובות, אך עדיין נכשל ב-27.5% מהמקרים.
תקציר מקורי באנגליתarXiv:2610.03080v1 Announce Type: cross Abstract: Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית