כתבה
arXiv cs.CL ·
DFAH-Bench: Benchmarking Observable Agent Instability in Financial Decision-Making
תקציר מקורי באנגליתarXiv:2607.20491v1 Announce Type: cross Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time. We introduce DFAH-Bench, a replay benchmark that measures observable behavioral instability in financial agent decision-making across three channels -- tool-call trajectories, evidence contacts, and decision concentration -- none of which require access to hidden reasoning text. Across 8,127 replay episodes spanning 10 models and 3 financial tasks, we find that outcome agreement alone is an incomplete stability signal: frontier models can agree on decisions 95% of the time while following the same tool path only 77% of the time -- an 18-percentage-point gap (95% CI: [0.14, 0.22]) that outcome-on
קרא במקור המקורי
arxiv.org
פתח כתבה מקורית