כתבה
arXiv cs.CL ·
MintEval: האם LLMs מimplement את הסטרטגיה המסחרית שאתה פנה אליה?
MintEval: Do LLMs Implement the Trading Strategy You Asked For? A Behavioural-Equivalence Benchmark for Natural-Language-to-Strategy Code
MintEval היא בנקאות לבדיקת LLMs' יכולת לimplement סטרטגיות מסחריות. הבנקאות נועדה לבדוק אם LLMs מimplement את הסטרטגיה המסחרית שאתה פנה אליה, ולא רק לייצר תזכירי מסחר.
תקציר מקורי באנגליתarXiv:2610.03080v1 Announce Type: cross Abstract: Large language models are moving from producing trading signals to writing the code that executes them. The failure mode of the second role is silent: generated code runs, a backtest plots, yet the risk logic that the trader described is not the logic being executed. Existing code benchmarks test functional correctness on unit tests and finance benchmarks test forecasting; neither measures whether an implementation behaves like the strategy that was asked for. We introduce MintEval, a benchmark in which reference strategies are generated programmatically from a library of composable building blocks, back-translated into colloquial trader instructions, and re-implemented by the model under test. Generated and reference programs are executed
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית